How Whissle measures up
Voice agents are usually sold on demos. A demo shows you the happy path. These are results on open, tool-using agent benchmarks that score whether the agent actually completed the job — judged against the resulting database or chart state rather than against the transcript. τ²-bench for retail and airline support, three healthcare-agent benchmarks for clinical conversation and health-record work, and our own suite for the conversation flows customers actually build.
We run the same tasks twice: once in text, once spoken end to end through our real speech pipeline. The voice number is substantially worse, and it is on this page with its root cause named. Every score carries the sample it was computed over, what was excluded, who did the judging, and which model produced it — ours as specifically as theirs, so you can see which two configurations are being put side by side.
All results measured on the production stack. Last updated 8 August 2026. Served from the measured snapshot committed in the site repository — the results store holds no published runs yet, so there is nothing newer to show. These are the same numbers, not a cached copy of a different set.
Benchmarks
Retail — order support
τ²-bench
Whether the agent completes real order-support jobs — returns, exchanges, address changes — using the benchmark's own tools, judged on the database state it leaves behind.
61.4%70 of 114 tasks95% CI 52.2–69.8%TextIndependent judge — τ²-bench database-state checkerRan on
claude-haiku-4-5
the cheapest and fastest model in its family
provider Anthropic
No setup-matched baseline yet — 5 published figures shown for context, under a different setup
1 run on recordFull resultsAirline — booking & changes
τ²-bench
Whether the agent can book, change and cancel flights against a live task database while holding to the airline's policy constraints.
56.0%28 of 50 tasks95% CI 42.3–68.8%TextIndependent judge — τ²-bench database-state checkerRan on
claude-haiku-4-5
the cheapest and fastest model in its family
provider Anthropic
No setup-matched baseline yet — 5 published figures shown for context, under a different setup
1 run on recordFull resultsRetail — order support
τ³-bench (half-duplex voice)
The identical retail tasks, spoken end to end through the real speech pipeline. The gap against the text run is the cost of hearing instead of reading.
20.0%1 of 5 tasks95% CI 3.6–62.4%VoicePreliminaryIndependent judge — τ²-bench database-state checkerRan on
claude-haiku-4-5
the cheapest and fastest model in its family
provider Anthropic
Spoken through
recognition the production speech configuration — a third-party engine chosen by language routing; the run does not record which
synthesis the production synthesis configuration; the run does not record which engine served it
Whissle metadata head — not in the path
No published comparator
1 run on recordFull resultsHeadache intake — designed flow, over voice
Whissle conversation-flow suite
Whether a designed clinical-intake flow survives a real caller: red flags raised, topics covered out of order, uncertainty, time pressure — and whether the agent ends the call itself.
70.0%7 of 10 tasks95% CI 39.7–89.2%VoicePreliminaryOur own judge — Whissle deterministic flow analyzerRan on
claude-haiku-4-5
the cheapest and fastest model in its family
provider Anthropic
Spoken through
recognition the production speech configuration — a third-party engine chosen by language routing; the run does not record which
synthesis the production synthesis configuration; the run does not record which engine served it
Whissle metadata head — not in the path
No published comparator
2 runs on recordFull resultsFlow-edit sensitivity
Whissle conversation-flow suite
Whether an edit made in the flow designer actually reaches a live conversation — and whether a staged draft correctly reaches none. Probed over the text channel, so it says nothing about the speech path.
100.0%8 of 8 tasks95% CI 67.6–100.0%TextPreliminaryOur own judge — Whissle deterministic flow analyzerRan on
claude-haiku-4-5
the cheapest and fastest model in its family
provider Anthropic
No published comparator
1 run on recordFull resultsDefault-flow coverage across agent types
Whissle conversation-flow suite
Whether every shipped agent type boots with a default flow that actually attaches and runs. Driven over the text channel — it is a wiring check, not a voice result.
100.0%15 of 15 tasks95% CI 79.6–100.0%TextPreliminaryOur own judge — Whissle deterministic flow analyzerRan on
claude-haiku-4-5
the cheapest and fastest model in its family
provider Anthropic
No published comparator
1 run on recordFull resultsDiagnostic consultation
AgentClinic
Whether the agent can reach a diagnosis through a conversation — asking the questions that discriminate between candidates, ordering the tests it needs, and committing to an answer within a fixed inference budget.
75.0%75 of 100 tasks95% CI 65.7–82.5%TextOur own judge — rubric jury routed through Whissle's own model APIRan on
claude-haiku-4-5
the cheapest and fastest model in its family
provider Anthropic
No published comparator
1 run on recordFull resultsElectronic health record — read and write
MedAgentBench
Whether the agent can operate a patient chart over a standard health-records API: read the right resource for a clinical question, and write a correct, conformant one back. Graded deterministically against chart state.
54.0%54 of 100 tasks95% CI 44.3–63.4%TextIndependent judge — the benchmark's own deterministic graderRan on
claude-haiku-4-5
the cheapest and fastest model in its family
provider Anthropic
No setup-matched baseline yet — 11 published figures shown for context, under a different setup
1 run on recordFull resultsDefault flows across six agent types
Whissle conversation-flow suite
Whether the flow a customer gets out of the box holds up on a real voice call against a caller who improvises — across support, intake, reception, scheduling, rentals and collections, with no per-domain tuning.
61.2%30 of 49 tasks95% CI 47.2–73.6%VoicePreliminary16 of 65 excluded (24.6%)Our own judge — Whissle deterministic flow analyzerRan on
claude-haiku-4-5
the cheapest and fastest model in its family
provider Anthropic
Spoken through
recognition the production speech configuration — a third-party engine chosen by language routing; the run does not record which
synthesis the production synthesis configuration; the run does not record which engine served it
Whissle metadata head — not in the path
No published comparator
1 run on recordFull results
Every published run
| Benchmark | Suite | Modality | Result | Ran on | Excluded | Comparison |
|---|---|---|---|---|---|---|
| Retail — order support | τ²-bench | Text | 61.4%70 of 114 tasks | claude-haiku-4-5the cheapest and fastest model in its familyprovider Anthropic | None excluded | No setup-matched baseline yet — 5 published figures shown for context, under a different setup |
| Airline — booking & changes | τ²-bench | Text | 56.0%28 of 50 tasks | claude-haiku-4-5the cheapest and fastest model in its familyprovider Anthropic | None excluded | No setup-matched baseline yet — 5 published figures shown for context, under a different setup |
| Retail — order supportPreliminary | τ³-bench (half-duplex voice) | Voice | 20.0%1 of 5 tasks | claude-haiku-4-5the cheapest and fastest model in its familyprovider Anthropic | None excluded | No published comparator |
| Headache intake — designed flow, over voicePreliminary | Whissle conversation-flow suite | Voice | 70.0%7 of 10 tasks | claude-haiku-4-5the cheapest and fastest model in its familyprovider Anthropic | None excluded | No published comparator |
| Flow-edit sensitivityPreliminary | Whissle conversation-flow suite | Text | 100.0%8 of 8 tasks | claude-haiku-4-5the cheapest and fastest model in its familyprovider Anthropic | None excluded | No published comparator |
| Default-flow coverage across agent typesPreliminary | Whissle conversation-flow suite | Text | 100.0%15 of 15 tasks | claude-haiku-4-5the cheapest and fastest model in its familyprovider Anthropic | None excluded | No published comparator |
| Diagnostic consultation | AgentClinic | Text | 75.0%75 of 100 tasks | claude-haiku-4-5the cheapest and fastest model in its familyprovider Anthropic | None excluded | No published comparator |
| Electronic health record — read and write | MedAgentBench | Text | 54.0%54 of 100 tasks | claude-haiku-4-5the cheapest and fastest model in its familyprovider Anthropic | None excluded | No setup-matched baseline yet — 11 published figures shown for context, under a different setup |
| Default flows across six agent typesPreliminary | Whissle conversation-flow suite | Voice | 61.2%30 of 49 tasks | claude-haiku-4-5the cheapest and fastest model in its familyprovider Anthropic | 16 of 65 excluded (24.6%) | No published comparator |
The same tasks, spoken. This is where we are weakest.
Voice runs the identical agent on the identical tasks, but every turn has to survive the full speech pipeline. It costs us roughly 3.1× the score. We publish it because a diagnosed weak number is worth more to you than a page of wins.
| Retail — order support | Pass1 | Sample size |
|---|---|---|
| Text | 61.4 | N=114 |
| Voicepreliminary | 20.0 | N=5 |
N=5, one trial each. Wide error bars — treat the voice figure as directional, not as a confidence interval.
What actually goes wrong — and it isn't the reasoning
Speech recognition on spelled-out identifiers — order IDs, postcodes and, in the latest run, the customer's own name.
find_user_id_by_name_zip → { first_name: "Youssef", last_name: "Rossi", zip: "19122" }
The customer is yusuf_rossi_9620. The postcode was heard correctly; the name alone was enough to miss.
- Tool calls round-tripped
- 16
- The agent reached for the right tools and persisted through failures.
- Judged agent errors
- 0
- Across all five sessions. These are outcome failures, not conduct failures.
- Sessions ending naturally
- 5 / 5
- Every task terminated on a normal user stop — nothing hung or crashed.
We build our own speech recognition, so this is ours to fix rather than a vendor's to explain.
One of the five failures is not a voice failure
One of the five failures is ours in text too: the agent counted 12 t-shirt variants where 2 were flagged unavailable, and reported 12 instead of 10. That task authenticated cleanly and made 10 correct tool calls. It would fail identically in text.
Attributing every voice loss to speech recognition would have hidden that. It is on the page because it is on the transcript.
Our own suite: does the flow you designed survive a real caller?
τ²-bench measures an agent. It does not measure the thing our customers actually build — a multi-step conversation flow with gates, branches and tools. So we wrote a suite for that: a simulated caller with a persona and a goal improvises through a live voice session, and a deterministic analyser then audits the running machine against the flow's own declared contract.
| Measure | Latest | Previous run |
|---|---|---|
| Task success — headache intake, voiceTen distinct caller personas: urgent red flags, skipped topics, out-of-order answers, chronic uncertainty, time pressure. One of the ten never established a session (infrastructure, not agent behaviour) and is counted as a miss. | 7 / 10N=10 | 5 / 10 on 5 Aug |
| Reached a clean closeOur strictest bar: the agent has to end the call itself, not trail off once the caller's goal is met. Three of the six misses succeeded at the task but never closed. | 4 / 10N=10 | 3 / 10 on 5 Aug |
| Seeded agent types that attach and drive a default flowEvery shipped agent type — dental, patient check-in, medication, collections, tutoring, car rental and the rest — boots with a flow that actually runs. | 15 / 15N=15 | — |
Headache-intake scenario set, run over voice. Ten scenarios, one session each. The previous column is the identical scenario set run two days earlier — same tasks, same analyser, so the movement is real and not a change of yardstick.
A published flow edit reached the live conversation in every case, and a staged draft reached it in none.
Every edit travels the exact API the flow designer uses — never a backdoor. Validate, stage as a draft, confirm the live conversation shows no trace of it, publish, then go hunting for it in a real conversation. A failure here would be a product bug: an edit you made and shipped that the caller never hears.
Probed over the text channel. It shows an edit reaching the live agent — not that it survives the speech path.
| Edit kind | What changed | Signal required in the live conversation | Result |
|---|---|---|---|
| Spoken line | Rewrite what a step says | New line appears in the agent's opening turn | Picked up |
| Goal | Change what a step must find out | Agent asks the new question | Picked up |
| Transition condition | Narrow when an edge fires | Routing holds, then fires only on the new condition | Picked up |
| Transition target | Point an edge at a different step | New step entered, old one never visited | Picked up |
| Tool gate — remove | Strip a tool from a step | Tool absent from the gate and never invoked | Picked up |
| Tool gate — add | Grant a tool to a tool-less step | Gate now admits it | Picked up |
| Remove a step | Delete a mid-flow step, rewire inbound edges | Session skips it and reaches the next step | Picked up |
| Set a variable | Insert a variable step plus an expression edge | Variable set in the trace, edge fires, routes to close | Picked up |
The caller was satisfied. Did the agent follow the flow?
Task success asks whether the caller got what they wanted. It says nothing about whether the agent got there the way the flow said to — an agent can satisfy someone while skipping a step, leaking a tool into a state that should not have had it, or looping, and still score a clean pass. If you designed the flow because the steps are the point, this is the number you wanted. Same sessions as above; a different question.
Clean sessions
63.3%
31 of 49 scored sessions had no high-severity divergence
Not one finding at all
34.7%
17 of 49 — the strict bar, published because the headline is the lenient one
Divergences per 100 turns
9.9
41 findings over 414 turns. Measured per turn, so ending a call early cannot improve it
How this is scored. The share of sessions in which a deterministic analyzer, replaying the call against the flow's own declared specification, found nothing of high severity. There is no weighting and nothing to tune: a session either had a high-severity divergence or it did not. The full per-type table is below, so a reader who disagrees with how we rank severity can recompute it rather than take our word for the ranking.
| Divergence | Count | Per 100 turns | Severity |
|---|---|---|---|
| Never closed a finished callagent_no_closeThe caller's goal was met and the agent kept the line open instead of ending the call. The single most common divergence, by a wide margin. | 16 | 3.9 | high |
| Ended before the flow was donepremature_terminationThe opposite failure: the conversation stopped with steps still outstanding. Both directions of the same missing judgement about when a call is over. | 9 | 2.2 | medium |
| Could not get to an end statestuck_terminationThe flow reached a step it could not leave, and the session ran out rather than finishing. | 7 | 1.7 | medium |
| Circled the same stepsstuck_loopThe machine re-entered states it had already visited without making progress — usually re-asking something the caller had already answered. | 7 | 1.7 | medium |
| Took a route the flow does not defineillegal_transitionThe conversation moved between steps by an edge that does not exist in the specification. Twice, in 414 turns. | 2 | 0.5 | high |
No new runs were needed for any of this: the analyzer was already recording its findings into every session, so these are recoverable for every run the suite has ever done. Said plainly because a new reading of existing data is a weaker claim than a fresh measurement, and the difference should be visible rather than inferred.
The state machine is not the problem
Two illegal transitions across 414 turns, and no tool leaked into a step that should not have had it, no variable fell out of sync with its trace, no guard was violated. The engine executes the flow it was given. Nearly every finding here is about something else.
Everything else is about ending the call
Thirty-nine of the forty-one findings are one of four things: not closing a call whose goal was met, closing one that was not finished, getting stuck at a step, or looping. Those are the same missing judgement seen from four angles — the agent does not reliably know when a conversation is over. It is also the weakness this page already reports from a different direction, where reaching a clean close was the strictest bar and the one most often missed.
A clean pass on the task is not a clean pass on the flow
Task success across the same sessions is 61.2%; adherence is 63.3%, and they are not the same 61 to 63 percent of sessions. An agent can satisfy a caller while diverging from the flow, and it can follow the flow faithfully to an outcome the caller did not want. Publishing one without the other would let either failure hide behind the other's number.
What this is not
It is one analyzer, written by us, auditing our own product against a specification we also wrote. It is deterministic rather than an opinion, and it is not third-party validation. Forty-nine sessions across six domains is enough to see the shape above and not enough to rank the domains against each other.
How much of the score is the agent, and how much is the model under it?
Every other number on this page was produced on one arm — whatever the production default was that week — which leaves the obvious question unanswered. So the two healthcare benchmarks were re-run across seven brains with everything else held still. This is the one experiment here where the model is the variable rather than a disclosure.
These runs were graded independently. Graded by an external provider's model, independent of the agent under test — which is a stronger footing than the larger healthcare runs elsewhere on this page, every one of which was graded through our own model API. Twenty-five cases per arm against those runs' hundred, so this buys independence at the price of sample size, and both facts are stated rather than traded off silently.
Accuracy, 25 cases per arm
| Model | Health record — read & write | Diagnosis | Declined to answer |
|---|---|---|---|
| claude-fable-55 sessions lost to infrastructure and excluded | 75.0%20 tasks | 96.0%25 cases | — |
| claude-opus-5 | 72.0%25 tasks | 92.0%25 cases | — |
| gemini-3.5-flash | 72.0%25 tasks | 84.0%25 cases | — |
| claude-haiku-4-5 | 68.0%25 tasks | 84.0%25 cases | 4.0% |
| gemini-3.5-flash-lite | 68.0%25 tasks | 92.0%25 cases | — |
| claude-sonnet-5 | 52.0%25 tasks | 88.0%25 cases | — |
| gemini-3-flash-preview | 52.0%25 tasks | 60.0%25 cases | — |
Twenty-five cases per arm. At that sample the gap between adjacent rows is inside the noise, so this table does not rank arms that sit a few points apart and is not offered as doing so. Sorted by health-record success, which is the deterministically graded half.
What each arm costs to serve
A separate experiment, and the two are not merged. This one sweeps reasoning-depth settings that the accuracy run did not vary, so putting both on one row per model would imply a joint measurement nobody made. Fifteen samples per arm on a fixed set of representative turns.
| Arm | Cost / 1,000 turns | Time to first token — median | 95th percentile |
|---|---|---|---|
| gemini-3.5-flash-lite, minimal depth | $0.20 | 441 ms | 506 ms |
| claude-haiku-4-5 (the production default) | $0.86 | 479 ms | 737 ms |
| gemini-3.5-flash, minimal depth | $0.93 | 722 ms | 817 ms |
| claude-sonnet-5, low depth, no thinking | $2.21 | 1191 ms | 1697 ms |
| gemini-3.5-flash, low depth | $2.25 | 1254 ms | 2149 ms |
| claude-sonnet-5, low depth | $2.36 | 1506 ms | 2965 ms |
| gemini-3-flash-preview (the failover default) | $3.04 | 1522 ms | 2516 ms |
| claude-sonnet-5, default depth | $3.22 | 1957 ms | 6621 ms |
| gemini-3.5-flash, default depth | $4.52 | 2054 ms | 3525 ms |
| claude-opus-5, low depth, no thinking | $5.29 | 779 ms | 1165 ms |
| claude-opus-5, low depth | $5.31 | 1552 ms | 2358 ms |
| claude-opus-5, default depth | $5.81 | 1605 ms | 2222 ms |
| claude-opus-5, medium depth | $6.37 | 1743 ms | 4052 ms |
| claude-fable-5, default depth | $15.23 | 3144 ms | 5765 ms |
claude-opus-5, low depth, fast serving
Every attempt errored and none returned a token, so the arm has no row above. It is named here rather than dropped: a reader comparing depth settings should know one of them could not be measured at all.
How to read this
The cheapest arm is not the worst one
The production default sits mid-table on health-record accuracy and matches a model roughly seven times its cost on diagnosis. The two lightest arms in the serving table cost $0.20 and $0.86 per thousand turns against $15.23 at the top, and neither is last on either accuracy column. Whatever the right arm is, 'spend more' is not a finding this run supports.
Depth settings move responsiveness more than they move cost
Turning thinking off on the largest model halves its time to first token — 1552ms to 779ms — and changes the bill by two cents per thousand turns. The two knobs are close to independent, which means the responsiveness that a voice call needs can be bought without moving to a cheaper brain.
The two halves disagree, and that is the useful part
Ranking by health-record success and ranking by diagnostic accuracy do not produce the same order — one arm is second on the first and joint-fourth on the second. Reading tools and charts correctly is a different skill from reasoning to an answer, and an agent that has to do both should not be selected on a single blended figure.
What twenty-five cases can and cannot settle
At this sample the gap between adjacent rows is inside the noise, so the table cannot rank arms that sit a few points apart and is not presented as doing so. What it does separate is the top from the bottom of each column, and the shape of the cost curve against them.
How these runs were configured
Everything that could quietly move a score is stated here. If a detail you need to judge the numbers is missing, it is an omission, not a decision — tell us and it goes on the page.
- What is under test
- The production Whissle agent stack — the same prompt, tool layer and guardrails a customer's agent runs on — driven over our public API. The benchmark owns the tools, the task database and the scoring; we only supply the agent.
- Which model produced the number
- Stated per run, next to the score, on the card and in the table — model id, and the provider, effort, thinking budget and fast-mode setting wherever the run pinned one. A run that does not record its arm says "not recorded" rather than inheriting whatever we deploy today. The comparison rows name their models in the same words, so a reader can see which two configurations are being placed side by side.
- Scoring basis
- Primarily database state: did the world end up the way the task required. Not a judge's opinion of the transcript, and not how fluent the agent sounded.
- User simulator
- On the τ² retail, airline and half-duplex voice runs: GPT-4o, the τ² default. In the voice runs its turns are spoken aloud by a third-party synthetic voice, so our speech recognition hears audio it has never been tuned on. The conversation-flow suite is different — its simulated caller and its judges route through Whissle's own model API without pinning a model, and those runs do not record which model served them. We are saying that rather than reusing the τ² answer for a suite it does not describe.
- Voice path
- Half-duplex — one turn at a time — over a real WebRTC session, through the full speech-understanding → agent → speech-synthesis cascade. Tool calls are delegated back to the benchmark so it stays the authority on tools and scoring.
- Concurrency
- 2. Higher concurrency triggers upstream rate limiting, which pollutes the score rather than measuring the agent.
- Trials
- Pass^1 — one attempt per task, no best-of-n, no retries at the harness level.
The harness is open. So are the trajectories.
The τ² and τ³ suites, our voice harness, both conversation-flow suites and the raw per-task trajectories behind every number on this page live in WhissleAI/tau2-bench-w. Point it at an agent you built on Whissle and you get your own numbers, not ours.
git clone https://github.com/WhissleAI/tau2-bench-w && cd tau2-bench-w
uv sync
export WHISSLE_BASE=... WHISSLE_AGENT_ID=... WHISSLE_API_KEY=...
# text — the two rows in the headline table
uv run tau2 run --domain retail --agent whissle --agent-llm whissle \
--user-llm gpt-4o --max-concurrency 2
uv run tau2 run --domain airline --agent whissle --agent-llm whissle \
--user-llm gpt-4o --max-concurrency 2
# voice — half-duplex, through the real speech pipeline
# args: <domain> <tasks> <concurrency> <max-steps>
./run_hd.sh retail 5 2 40
# conversation-flow suites
./run_flow_sim.sh --agent-type headache_enrollment --sessions 10
./run_flow_mutation.sh run --agent-type headache_enrollmentWhere each number on this page comes from
- Retail — order support · text
results/whissle/retail_run1.json - Airline — booking & changes · text
results/whissle/airline_run1.json - Retail — order support · voice
results/whissle/hd_retail_n5.json - Flow-edit sensitivity
results/whissle/flow_mutation/headache_enrollment/REPORT.md - Conversation-flow suite
results/whissle/flow_sim/
Things we tried that didn't work
A benchmark page that only contains wins is a marketing asset. These are on the page for the same reason the voice score is.
A prompt change we could have claimed as a win, and didn't
We added an instruction telling the agent to complete actions rather than promise them, persist through authentication, and treat handing off to a human as a last resort. Retail moved 61.4 → 62.3 (70 → 71 of 114). That is one task, inside run-to-run noise, and the number of handoffs to a human was effectively unchanged. We kept the instruction because it is good behaviour, and we claim no lift from it.
A harness fix that changed nothing
We found and fixed a bug that triplicated transcript lines, then re-ran the voice retail suite expecting a bump. Pass^1 came back at 20.0 — identical. It should have: scoring is mostly database state, and the bug only polluted what the simulated caller read back. We are reporting that plainly rather than implying a fix that did not move the headline.
Our own harness was scoring our billing failures as agent failures
The conversation-flow suite ran 65 sessions across six agent types. 16 of them never produced a single word: 13 because the SIMULATED CALLER'S model calls were refused for insufficient credit on our own workspace, and 3 because the session's data channel died. Two domains had those correctly excluded. Two did not — so debt collection was published at 9.1% when 7 of its 11 sessions were a billing error and only 4 were conversations, and car rental at 45.5% when 6 of 11 never started. Correcting it moves the suite from 46.2% to 61.2%, which is a large move in our favour off our own bug report, so the page prints both numbers and the floor is the one to argue with. The harness fix — treat a payment error on the simulator endpoint as infrastructure, exactly as the voice-channel error already is — is filed against the public repo.
Below the published leaderboard on health records, printed anyway
On the health-record benchmark the agent scored 54.0% against published figures of 69.7, 64.0 and 62.7 for three widely-used models. It is the most directly comparable number on this page — same task grammar, same deterministic grader, no rubric and no judge model on either side — which is exactly why it is not buried. Two caveats bound it in our favour and are stated on the card rather than used as an excuse: we scored 100 of the 300 published tasks, and the sandbox accepted 7 non-conformant writes that a real records system would have rejected.
Voice error bars are wide and we are not hiding behind them
The voice number is five tasks, one trial each, all with the same simulated customer — so the name-recognition failure is one sample repeated, not four independent ones. It is directional. A broader run across distinct names and identifiers is the next thing we owe this page.
The voice gap, turned into work
Every item below points at a specific failure in the runs above. The next voice numbers land on this page whichever way they go.
- 01
Retry before you hand off
In the failing runs the agent gave up after a single failed lookup. One task re-asked for the name, heard it correctly the second time, and went on to make ten correct tool calls — so the recognition ceiling is not hard. Retrying an identifier once, and reading it back digit by digit, is a cheaper fix than a better model.
- 02
Prefer identifiers a machine can check
Where a flow can authenticate on something validatable against a known set rather than on a spelled-out string, it should. That removes the failure mode instead of mitigating it.
- 03
Tune recognition for alphanumerics
Order IDs, postcodes and names are where the losses are concentrated. We own the speech recognition, so this is a model and decoding problem we can go and solve rather than a vendor ticket.
- 04
Widen the voice sample
Five tasks is enough to diagnose and not enough to claim. Broader voice runs across distinct customers and both domains are queued, and they will land on this page whichever way they go.