How Whissle measures up

Voice agents are usually sold on demos. A demo shows you the happy path. These are results on open, tool-using agent benchmarks that score whether the agent actually completed the job — judged against the resulting database or chart state rather than against the transcript. τ²-bench for retail and airline support, three healthcare-agent benchmarks for clinical conversation and health-record work, and our own suite for the conversation flows customers actually build.

We run the same tasks twice: once in text, once spoken end to end through our real speech pipeline. The voice number is substantially worse, and it is on this page with its root cause named. Every score carries the sample it was computed over, what was excluded, who did the judging, and which model produced it — ours as specifically as theirs, so you can see which two configurations are being put side by side.

All results measured on the production stack. Last updated 8 August 2026. Served from the measured snapshot committed in the site repository — the results store holds no published runs yet, so there is nothing newer to show. These are the same numbers, not a cached copy of a different set.

Benchmarks

  • Whether the agent completes real order-support jobs — returns, exchanges, address changes — using the benchmark's own tools, judged on the database state it leaves behind.

    61.4%70 of 114 tasks
    95% CI 52.2–69.8%
    TextIndependent judge — τ²-bench database-state checker

    Ran on

    claude-haiku-4-5

    the cheapest and fastest model in its family

    provider Anthropic

    No setup-matched baseline yet — 5 published figures shown for context, under a different setup

    1 run on recordFull results
  • Whether the agent can book, change and cancel flights against a live task database while holding to the airline's policy constraints.

    56.0%28 of 50 tasks
    95% CI 42.3–68.8%
    TextIndependent judge — τ²-bench database-state checker

    Ran on

    claude-haiku-4-5

    the cheapest and fastest model in its family

    provider Anthropic

    No setup-matched baseline yet — 5 published figures shown for context, under a different setup

    1 run on recordFull results
  • Retail — order support

    τ³-bench (half-duplex voice)

    The identical retail tasks, spoken end to end through the real speech pipeline. The gap against the text run is the cost of hearing instead of reading.

    20.0%1 of 5 tasks
    95% CI 3.6–62.4%
    VoicePreliminaryIndependent judge — τ²-bench database-state checker

    Ran on

    claude-haiku-4-5

    the cheapest and fastest model in its family

    provider Anthropic

    Spoken through

    recognition the production speech configuration — a third-party engine chosen by language routing; the run does not record which

    synthesis the production synthesis configuration; the run does not record which engine served it

    Whissle metadata head — not in the path

    No published comparator

    1 run on recordFull results
  • Headache intake — designed flow, over voice

    Whissle conversation-flow suite

    Whether a designed clinical-intake flow survives a real caller: red flags raised, topics covered out of order, uncertainty, time pressure — and whether the agent ends the call itself.

    70.0%7 of 10 tasks
    95% CI 39.7–89.2%
    VoicePreliminaryOur own judge — Whissle deterministic flow analyzer

    Ran on

    claude-haiku-4-5

    the cheapest and fastest model in its family

    provider Anthropic

    Spoken through

    recognition the production speech configuration — a third-party engine chosen by language routing; the run does not record which

    synthesis the production synthesis configuration; the run does not record which engine served it

    Whissle metadata head — not in the path

    No published comparator

    2 runs on recordFull results
  • Flow-edit sensitivity

    Whissle conversation-flow suite

    Whether an edit made in the flow designer actually reaches a live conversation — and whether a staged draft correctly reaches none. Probed over the text channel, so it says nothing about the speech path.

    100.0%8 of 8 tasks
    95% CI 67.6–100.0%
    TextPreliminaryOur own judge — Whissle deterministic flow analyzer

    Ran on

    claude-haiku-4-5

    the cheapest and fastest model in its family

    provider Anthropic

    No published comparator

    1 run on recordFull results
  • Default-flow coverage across agent types

    Whissle conversation-flow suite

    Whether every shipped agent type boots with a default flow that actually attaches and runs. Driven over the text channel — it is a wiring check, not a voice result.

    100.0%15 of 15 tasks
    95% CI 79.6–100.0%
    TextPreliminaryOur own judge — Whissle deterministic flow analyzer

    Ran on

    claude-haiku-4-5

    the cheapest and fastest model in its family

    provider Anthropic

    No published comparator

    1 run on recordFull results
  • Whether the agent can reach a diagnosis through a conversation — asking the questions that discriminate between candidates, ordering the tests it needs, and committing to an answer within a fixed inference budget.

    75.0%75 of 100 tasks
    95% CI 65.7–82.5%
    TextOur own judge — rubric jury routed through Whissle's own model API

    Ran on

    claude-haiku-4-5

    the cheapest and fastest model in its family

    provider Anthropic

    No published comparator

    1 run on recordFull results
  • Whether the agent can operate a patient chart over a standard health-records API: read the right resource for a clinical question, and write a correct, conformant one back. Graded deterministically against chart state.

    54.0%54 of 100 tasks
    95% CI 44.3–63.4%
    TextIndependent judge — the benchmark's own deterministic grader

    Ran on

    claude-haiku-4-5

    the cheapest and fastest model in its family

    provider Anthropic

    No setup-matched baseline yet — 11 published figures shown for context, under a different setup

    1 run on recordFull results
  • Default flows across six agent types

    Whissle conversation-flow suite

    Whether the flow a customer gets out of the box holds up on a real voice call against a caller who improvises — across support, intake, reception, scheduling, rentals and collections, with no per-domain tuning.

    61.2%30 of 49 tasks
    95% CI 47.2–73.6%
    VoicePreliminary16 of 65 excluded (24.6%)Our own judge — Whissle deterministic flow analyzer

    Ran on

    claude-haiku-4-5

    the cheapest and fastest model in its family

    provider Anthropic

    Spoken through

    recognition the production speech configuration — a third-party engine chosen by language routing; the run does not record which

    synthesis the production synthesis configuration; the run does not record which engine served it

    Whissle metadata head — not in the path

    No published comparator

    1 run on recordFull results
All results

Every published run

Latest run per benchmark. The sample size and the model that produced it sit next to every score, and anything preliminary or partially excluded says so in the same row rather than in a note underneath.
Latest published run for each Whissle benchmark, with suite, modality, sample size, score, the model configuration that produced it, exclusions and comparison status.
BenchmarkSuiteModalityResultRan onExcludedComparison
Retail — order supportτ²-benchText61.4%70 of 114 tasksclaude-haiku-4-5the cheapest and fastest model in its familyprovider AnthropicNone excludedNo setup-matched baseline yet — 5 published figures shown for context, under a different setup
Airline — booking & changesτ²-benchText56.0%28 of 50 tasksclaude-haiku-4-5the cheapest and fastest model in its familyprovider AnthropicNone excludedNo setup-matched baseline yet — 5 published figures shown for context, under a different setup
Retail — order supportPreliminaryτ³-bench (half-duplex voice)Voice20.0%1 of 5 tasksclaude-haiku-4-5the cheapest and fastest model in its familyprovider AnthropicNone excludedNo published comparator
Headache intake — designed flow, over voicePreliminaryWhissle conversation-flow suiteVoice70.0%7 of 10 tasksclaude-haiku-4-5the cheapest and fastest model in its familyprovider AnthropicNone excludedNo published comparator
Flow-edit sensitivityPreliminaryWhissle conversation-flow suiteText100.0%8 of 8 tasksclaude-haiku-4-5the cheapest and fastest model in its familyprovider AnthropicNone excludedNo published comparator
Default-flow coverage across agent typesPreliminaryWhissle conversation-flow suiteText100.0%15 of 15 tasksclaude-haiku-4-5the cheapest and fastest model in its familyprovider AnthropicNone excludedNo published comparator
Diagnostic consultationAgentClinicText75.0%75 of 100 tasksclaude-haiku-4-5the cheapest and fastest model in its familyprovider AnthropicNone excludedNo published comparator
Electronic health record — read and writeMedAgentBenchText54.0%54 of 100 tasksclaude-haiku-4-5the cheapest and fastest model in its familyprovider AnthropicNone excludedNo setup-matched baseline yet — 11 published figures shown for context, under a different setup
Default flows across six agent typesPreliminaryWhissle conversation-flow suiteVoice61.2%30 of 49 tasksclaude-haiku-4-5the cheapest and fastest model in its familyprovider Anthropic16 of 65 excluded (24.6%)No published comparator
Text vs voice

The same tasks, spoken. This is where we are weakest.

Voice runs the identical agent on the identical tasks, but every turn has to survive the full speech pipeline. It costs us roughly 3.1× the score. We publish it because a diagnosed weak number is worth more to you than a page of wins.

Retail order support: Pass^1 in text compared with Pass^1 in voice.
Retail — order supportPass1Sample size
Text61.4N=114
Voicepreliminary20.0N=5

N=5, one trial each. Wide error bars — treat the voice figure as directional, not as a confidence interval.

What actually goes wrong — and it isn't the reasoning

Speech recognition on spelled-out identifiers — order IDs, postcodes and, in the latest run, the customer's own name.

find_user_id_by_name_zip → { first_name: "Youssef", last_name: "Rossi", zip: "19122" }

The customer is yusuf_rossi_9620. The postcode was heard correctly; the name alone was enough to miss.

Tool calls round-tripped
16
The agent reached for the right tools and persisted through failures.
Judged agent errors
0
Across all five sessions. These are outcome failures, not conduct failures.
Sessions ending naturally
5 / 5
Every task terminated on a normal user stop — nothing hung or crashed.

We build our own speech recognition, so this is ours to fix rather than a vendor's to explain.

One of the five failures is not a voice failure

One of the five failures is ours in text too: the agent counted 12 t-shirt variants where 2 were flagged unavailable, and reported 12 instead of 10. That task authenticated cleanly and made 10 correct tool calls. It would fail identically in text.

Attributing every voice loss to speech recognition would have hidden that. It is on the page because it is on the transcript.

Conversation flows

Our own suite: does the flow you designed survive a real caller?

τ²-bench measures an agent. It does not measure the thing our customers actually build — a multi-step conversation flow with gates, branches and tools. So we wrote a suite for that: a simulated caller with a persona and a goal improvises through a live voice session, and a deterministic analyser then audits the running machine against the flow's own declared contract.

Conversation-flow suite results, with sample size and the previous run on the identical scenario set.
MeasureLatestPrevious run
Task success — headache intake, voiceTen distinct caller personas: urgent red flags, skipped topics, out-of-order answers, chronic uncertainty, time pressure. One of the ten never established a session (infrastructure, not agent behaviour) and is counted as a miss.7 / 10N=105 / 10 on 5 Aug
Reached a clean closeOur strictest bar: the agent has to end the call itself, not trail off once the caller's goal is met. Three of the six misses succeeded at the task but never closed.4 / 10N=103 / 10 on 5 Aug
Seeded agent types that attach and drive a default flowEvery shipped agent type — dental, patient check-in, medication, collections, tutoring, car rental and the rest — boots with a flow that actually runs.15 / 15N=15

Headache-intake scenario set, run over voice. Ten scenarios, one session each. The previous column is the identical scenario set run two days earlier — same tasks, same analyser, so the movement is real and not a change of yardstick.

Flow-edit sensitivity8 / 8

A published flow edit reached the live conversation in every case, and a staged draft reached it in none.

Every edit travels the exact API the flow designer uses — never a backdoor. Validate, stage as a draft, confirm the live conversation shows no trace of it, publish, then go hunting for it in a real conversation. A failure here would be a product bug: an edit you made and shipped that the caller never hears.

Probed over the text channel. It shows an edit reaching the live agent — not that it survives the speech path.

The eight flow edits exercised, what each changed, the behaviour the live conversation had to show, and the verdict.
Edit kindWhat changedSignal required in the live conversationResult
Spoken lineRewrite what a step saysNew line appears in the agent's opening turnPicked up
GoalChange what a step must find outAgent asks the new questionPicked up
Transition conditionNarrow when an edge firesRouting holds, then fires only on the new conditionPicked up
Transition targetPoint an edge at a different stepNew step entered, old one never visitedPicked up
Tool gate — removeStrip a tool from a stepTool absent from the gate and never invokedPicked up
Tool gate — addGrant a tool to a tool-less stepGate now admits itPicked up
Remove a stepDelete a mid-flow step, rewire inbound edgesSession skips it and reaches the next stepPicked up
Set a variableInsert a variable step plus an expression edgeVariable set in the trace, edge fires, routes to closePicked up
Flow adherence

The caller was satisfied. Did the agent follow the flow?

Task success asks whether the caller got what they wanted. It says nothing about whether the agent got there the way the flow said to — an agent can satisfy someone while skipping a step, leaking a tool into a state that should not have had it, or looping, and still score a clean pass. If you designed the flow because the steps are the point, this is the number you wanted. Same sessions as above; a different question.

Clean sessions

63.3%

31 of 49 scored sessions had no high-severity divergence

Not one finding at all

34.7%

17 of 49 — the strict bar, published because the headline is the lenient one

Divergences per 100 turns

9.9

41 findings over 414 turns. Measured per turn, so ending a call early cannot improve it

How this is scored. The share of sessions in which a deterministic analyzer, replaying the call against the flow's own declared specification, found nothing of high severity. There is no weighting and nothing to tune: a session either had a high-severity divergence or it did not. The full per-type table is below, so a reader who disagrees with how we rank severity can recompute it rather than take our word for the ranking.

Divergences between the running conversation and its declared flow, by type, with counts and per-100-turn density.
DivergenceCountPer 100 turnsSeverity
Never closed a finished callagent_no_closeThe caller's goal was met and the agent kept the line open instead of ending the call. The single most common divergence, by a wide margin.163.9high
Ended before the flow was donepremature_terminationThe opposite failure: the conversation stopped with steps still outstanding. Both directions of the same missing judgement about when a call is over.92.2medium
Could not get to an end statestuck_terminationThe flow reached a step it could not leave, and the session ran out rather than finishing.71.7medium
Circled the same stepsstuck_loopThe machine re-entered states it had already visited without making progress — usually re-asking something the caller had already answered.71.7medium
Took a route the flow does not defineillegal_transitionThe conversation moved between steps by an edge that does not exist in the specification. Twice, in 414 turns.20.5high

No new runs were needed for any of this: the analyzer was already recording its findings into every session, so these are recoverable for every run the suite has ever done. Said plainly because a new reading of existing data is a weaker claim than a fresh measurement, and the difference should be visible rather than inferred.

The state machine is not the problem

Two illegal transitions across 414 turns, and no tool leaked into a step that should not have had it, no variable fell out of sync with its trace, no guard was violated. The engine executes the flow it was given. Nearly every finding here is about something else.

Everything else is about ending the call

Thirty-nine of the forty-one findings are one of four things: not closing a call whose goal was met, closing one that was not finished, getting stuck at a step, or looping. Those are the same missing judgement seen from four angles — the agent does not reliably know when a conversation is over. It is also the weakness this page already reports from a different direction, where reaching a clean close was the strictest bar and the one most often missed.

A clean pass on the task is not a clean pass on the flow

Task success across the same sessions is 61.2%; adherence is 63.3%, and they are not the same 61 to 63 percent of sessions. An agent can satisfy a caller while diverging from the flow, and it can follow the flow faithfully to an outcome the caller did not want. Publishing one without the other would let either failure hide behind the other's number.

What this is not

It is one analyzer, written by us, auditing our own product against a specification we also wrote. It is deterministic rather than an opinion, and it is not third-party validation. Forty-nine sessions across six domains is enough to see the shape above and not enough to rank the domains against each other.

Which brain

How much of the score is the agent, and how much is the model under it?

Every other number on this page was produced on one arm — whatever the production default was that week — which leaves the obvious question unanswered. So the two healthcare benchmarks were re-run across seven brains with everything else held still. This is the one experiment here where the model is the variable rather than a disclosure.

These runs were graded independently. Graded by an external provider's model, independent of the agent under test — which is a stronger footing than the larger healthcare runs elsewhere on this page, every one of which was graded through our own model API. Twenty-five cases per arm against those runs' hundred, so this buys independence at the price of sample size, and both facts are stated rather than traded off silently.

Accuracy, 25 cases per arm

Health-record and diagnostic accuracy for seven models, with the scored denominator on every figure.
ModelHealth record — read & writeDiagnosisDeclined to answer
claude-fable-55 sessions lost to infrastructure and excluded75.0%20 tasks96.0%25 cases
claude-opus-572.0%25 tasks92.0%25 cases
gemini-3.5-flash72.0%25 tasks84.0%25 cases
claude-haiku-4-568.0%25 tasks84.0%25 cases4.0%
gemini-3.5-flash-lite68.0%25 tasks92.0%25 cases
claude-sonnet-552.0%25 tasks88.0%25 cases
gemini-3-flash-preview52.0%25 tasks60.0%25 cases

Twenty-five cases per arm. At that sample the gap between adjacent rows is inside the noise, so this table does not rank arms that sit a few points apart and is not offered as doing so. Sorted by health-record success, which is the deterministically graded half.

What each arm costs to serve

A separate experiment, and the two are not merged. This one sweeps reasoning-depth settings that the accuracy run did not vary, so putting both on one row per model would imply a joint measurement nobody made. Fifteen samples per arm on a fixed set of representative turns.

Serving cost per thousand turns and time to first token for fourteen model configurations.
ArmCost / 1,000 turnsTime to first token — median95th percentile
gemini-3.5-flash-lite, minimal depth$0.20441 ms506 ms
claude-haiku-4-5 (the production default)$0.86479 ms737 ms
gemini-3.5-flash, minimal depth$0.93722 ms817 ms
claude-sonnet-5, low depth, no thinking$2.211191 ms1697 ms
gemini-3.5-flash, low depth$2.251254 ms2149 ms
claude-sonnet-5, low depth$2.361506 ms2965 ms
gemini-3-flash-preview (the failover default)$3.041522 ms2516 ms
claude-sonnet-5, default depth$3.221957 ms6621 ms
gemini-3.5-flash, default depth$4.522054 ms3525 ms
claude-opus-5, low depth, no thinking$5.29779 ms1165 ms
claude-opus-5, low depth$5.311552 ms2358 ms
claude-opus-5, default depth$5.811605 ms2222 ms
claude-opus-5, medium depth$6.371743 ms4052 ms
claude-fable-5, default depth$15.233144 ms5765 ms
Arm that produced nothing

claude-opus-5, low depth, fast serving

Every attempt errored and none returned a token, so the arm has no row above. It is named here rather than dropped: a reader comparing depth settings should know one of them could not be measured at all.

How to read this

The cheapest arm is not the worst one

The production default sits mid-table on health-record accuracy and matches a model roughly seven times its cost on diagnosis. The two lightest arms in the serving table cost $0.20 and $0.86 per thousand turns against $15.23 at the top, and neither is last on either accuracy column. Whatever the right arm is, 'spend more' is not a finding this run supports.

Depth settings move responsiveness more than they move cost

Turning thinking off on the largest model halves its time to first token — 1552ms to 779ms — and changes the bill by two cents per thousand turns. The two knobs are close to independent, which means the responsiveness that a voice call needs can be bought without moving to a cheaper brain.

The two halves disagree, and that is the useful part

Ranking by health-record success and ranking by diagnostic accuracy do not produce the same order — one arm is second on the first and joint-fourth on the second. Reading tools and charts correctly is a different skill from reasoning to an answer, and an agent that has to do both should not be selected on a single blended figure.

What twenty-five cases can and cannot settle

At this sample the gap between adjacent rows is inside the noise, so the table cannot rank arms that sit a few points apart and is not presented as doing so. What it does separate is the top from the bottom of each column, and the shape of the cost curve against them.

Methodology

How these runs were configured

Everything that could quietly move a score is stated here. If a detail you need to judge the numbers is missing, it is an omission, not a decision — tell us and it goes on the page.

What is under test
The production Whissle agent stack — the same prompt, tool layer and guardrails a customer's agent runs on — driven over our public API. The benchmark owns the tools, the task database and the scoring; we only supply the agent.
Which model produced the number
Stated per run, next to the score, on the card and in the table — model id, and the provider, effort, thinking budget and fast-mode setting wherever the run pinned one. A run that does not record its arm says "not recorded" rather than inheriting whatever we deploy today. The comparison rows name their models in the same words, so a reader can see which two configurations are being placed side by side.
Scoring basis
Primarily database state: did the world end up the way the task required. Not a judge's opinion of the transcript, and not how fluent the agent sounded.
User simulator
On the τ² retail, airline and half-duplex voice runs: GPT-4o, the τ² default. In the voice runs its turns are spoken aloud by a third-party synthetic voice, so our speech recognition hears audio it has never been tuned on. The conversation-flow suite is different — its simulated caller and its judges route through Whissle's own model API without pinning a model, and those runs do not record which model served them. We are saying that rather than reusing the τ² answer for a suite it does not describe.
Voice path
Half-duplex — one turn at a time — over a real WebRTC session, through the full speech-understanding → agent → speech-synthesis cascade. Tool calls are delegated back to the benchmark so it stays the authority on tools and scoring.
Concurrency
2. Higher concurrency triggers upstream rate limiting, which pollutes the score rather than measuring the agent.
Trials
Pass^1 — one attempt per task, no best-of-n, no retries at the harness level.
Reproduce it

The harness is open. So are the trajectories.

The τ² and τ³ suites, our voice harness, both conversation-flow suites and the raw per-task trajectories behind every number on this page live in WhissleAI/tau2-bench-w. Point it at an agent you built on Whissle and you get your own numbers, not ours.

git clone https://github.com/WhissleAI/tau2-bench-w && cd tau2-bench-w
uv sync
export WHISSLE_BASE=... WHISSLE_AGENT_ID=... WHISSLE_API_KEY=...

# text — the two rows in the headline table
uv run tau2 run --domain retail  --agent whissle --agent-llm whissle \
  --user-llm gpt-4o --max-concurrency 2
uv run tau2 run --domain airline --agent whissle --agent-llm whissle \
  --user-llm gpt-4o --max-concurrency 2

# voice — half-duplex, through the real speech pipeline
#   args: <domain> <tasks> <concurrency> <max-steps>
./run_hd.sh retail 5 2 40

# conversation-flow suites
./run_flow_sim.sh      --agent-type headache_enrollment --sessions 10
./run_flow_mutation.sh run --agent-type headache_enrollment

Where each number on this page comes from

  • Retail — order support · textresults/whissle/retail_run1.json
  • Airline — booking & changes · textresults/whissle/airline_run1.json
  • Retail — order support · voiceresults/whissle/hd_retail_n5.json
  • Flow-edit sensitivityresults/whissle/flow_mutation/headache_enrollment/REPORT.md
  • Conversation-flow suiteresults/whissle/flow_sim/
Results that went against us

Things we tried that didn't work

A benchmark page that only contains wins is a marketing asset. These are on the page for the same reason the voice score is.

A prompt change we could have claimed as a win, and didn't

We added an instruction telling the agent to complete actions rather than promise them, persist through authentication, and treat handing off to a human as a last resort. Retail moved 61.4 → 62.3 (70 → 71 of 114). That is one task, inside run-to-run noise, and the number of handoffs to a human was effectively unchanged. We kept the instruction because it is good behaviour, and we claim no lift from it.

A harness fix that changed nothing

We found and fixed a bug that triplicated transcript lines, then re-ran the voice retail suite expecting a bump. Pass^1 came back at 20.0 — identical. It should have: scoring is mostly database state, and the bug only polluted what the simulated caller read back. We are reporting that plainly rather than implying a fix that did not move the headline.

Our own harness was scoring our billing failures as agent failures

The conversation-flow suite ran 65 sessions across six agent types. 16 of them never produced a single word: 13 because the SIMULATED CALLER'S model calls were refused for insufficient credit on our own workspace, and 3 because the session's data channel died. Two domains had those correctly excluded. Two did not — so debt collection was published at 9.1% when 7 of its 11 sessions were a billing error and only 4 were conversations, and car rental at 45.5% when 6 of 11 never started. Correcting it moves the suite from 46.2% to 61.2%, which is a large move in our favour off our own bug report, so the page prints both numbers and the floor is the one to argue with. The harness fix — treat a payment error on the simulator endpoint as infrastructure, exactly as the voice-channel error already is — is filed against the public repo.

Below the published leaderboard on health records, printed anyway

On the health-record benchmark the agent scored 54.0% against published figures of 69.7, 64.0 and 62.7 for three widely-used models. It is the most directly comparable number on this page — same task grammar, same deterministic grader, no rubric and no judge model on either side — which is exactly why it is not buried. Two caveats bound it in our favour and are stated on the card rather than used as an excuse: we scored 100 of the 300 published tasks, and the sandbox accepted 7 non-conformant writes that a real records system would have rejected.

Voice error bars are wide and we are not hiding behind them

The voice number is five tasks, one trial each, all with the same simulated customer — so the name-recognition failure is one sample repeated, not four independent ones. It is directional. A broader run across distinct names and identifiers is the next thing we owe this page.

What we're fixing next

The voice gap, turned into work

Every item below points at a specific failure in the runs above. The next voice numbers land on this page whichever way they go.

  1. 01

    Retry before you hand off

    In the failing runs the agent gave up after a single failed lookup. One task re-asked for the name, heard it correctly the second time, and went on to make ten correct tool calls — so the recognition ceiling is not hard. Retrying an identifier once, and reading it back digit by digit, is a cheaper fix than a better model.

  2. 02

    Prefer identifiers a machine can check

    Where a flow can authenticate on something validatable against a known set rather than on a spelled-out string, it should. That removes the failure mode instead of mitigating it.

  3. 03

    Tune recognition for alphanumerics

    Order IDs, postcodes and names are where the losses are concentrated. We own the speech recognition, so this is a model and decoding problem we can go and solve rather than a vendor ticket.

  4. 04

    Widen the voice sample

    Five tasks is enough to diagnose and not enough to claim. Broader voice runs across distinct customers and both domains are queued, and they will land on this page whichever way they go.