τ²-bench
Retail — order support
Whether the agent completes real order-support jobs — returns, exchanges, address changes — using the benchmark's own tools, judged on the database state it leaves behind.
Ran on
claude-haiku-4-5
the cheapest and fastest model in its family
provider Anthropic
The harness pinned nothing — no model, no provider, no reasoning-depth or speed setting. Those request controls did not exist on the benchmark endpoint when these runs were made, so every arm inherited the production deployment default.
Read from the deployment default in force on each run date, not from the run's own record: the endpoint did not report which model served a turn until after these runs, and no run artifact names one. Automatic failover to the secondary provider (Gemini 2.5 Flash) is logged but not recorded per turn, so we cannot rule out that a small fraction of turns was served by it.
No setup-matched baseline yet — 5 published figures shown for context, under a different setup
Run 31 July 2026.
Every metric this run produced
| Metric | Value | Sample | Interval |
|---|---|---|---|
| Pass^1Headline95% Wilson score interval, computed here from the pass count. The harness does not emit an interval; this is derived, not reported. | 61.4% | 70 of 114 tasks | 95% CI 52.2–69.8% |
Run this yourself
You will need: Node 18+ and uv. A Whissle workspace key (Settings → API keys on whissle.ai) and an agent id, plus a key for the benchmark's own user simulator — that simulator is τ²-bench's, not ours, and the harness calls it directly.
git clone https://github.com/WhissleAI/tau2-bench-w && cd tau2-bench-w
uv sync
export WHISSLE_BASE=... WHISSLE_AGENT_ID=... WHISSLE_API_KEY=...
uv run tau2 run --domain retail --agent whissle --agent-llm whissle \
--user-llm gpt-4o --max-concurrency 2Concurrency 2 is not a default worth changing — higher triggers upstream rate limiting, which pollutes the score rather than measuring the agent.
Running a benchmark straight from the Whissle CLI is on the roadmap and does not exist yet — the CLI manages agents, calls and records today, and we are not going to print a command on this page that would fail when you paste it. Until it lands, the harness above is the path.
What we measured against
Our arm in every row below
Whissle ran on
claude-haiku-4-5
the cheapest and fastest model in its family
provider Anthropic
The harness pinned nothing — no model, no provider, no reasoning-depth or speed setting. Those request controls did not exist on the benchmark endpoint when these runs were made, so every arm inherited the production deployment default.
Read from the deployment default in force on each run date, not from the run's own record: the endpoint did not report which model served a turn until after these runs, and no run artifact names one. Automatic failover to the secondary provider (Gemini 2.5 Flash) is logged but not recorded per turn, so we cannot rule out that a small fraction of turns was served by it.
Scoring 61.4% over 114 scored tasks. Where a row below names a different model, that difference is listed with the others.
No setup-matched baseline yet
We have not yet run another model through this exact harness — same domain, same task set, same user simulator, same pass^1, same concurrency — so there is no head-to-head to show. Published figures for other systems appear below for context, greyed, with the reasons they are not directly comparable. When a matched run lands it will appear here as a real comparison.
Published figures, for context
These were produced by other people under other conditions. They are not subtracted from our number and no margin is claimed against them. The public τ² ledgers describe themselves as a record of what was reported rather than a controlled cross-provider ranking, and the rows below differ from ours on the dimensions listed.
| System | Published score | Why it is not a head-to-head | Source |
|---|---|---|---|
| GPT-4oPublished — we quoted it | 60.0% | Same domain and task set as ours, but the scaffold, user simulator and concurrency behind the published figure are not stated — so it is context, not a head-to-head. The ledger describes itself as a source ledger, not a controlled cross-provider ranking, and instructs readers to match domain, task release, agent model, user-simulator model, scaffold, prompts, trial count and pass^k before comparing rows.
| Public τ²-bench source ledger (benchlm.ai, accessed 8 August 2026) |
| Claude-3.5-SonnetPublished — we quoted it | 69.0% | Same domain and task set as ours, but the scaffold, user simulator and concurrency behind the published figure are not stated — so it is context, not a head-to-head. The ledger describes itself as a source ledger, not a controlled cross-provider ranking, and instructs readers to match domain, task release, agent model, user-simulator model, scaffold, prompts, trial count and pass^k before comparing rows.
| Public τ²-bench source ledger (benchlm.ai, accessed 8 August 2026) |
| GLM-5.2Published — we quoted itat the top of the public ledger | 99.1% | Different setup — not directly comparable. Measured on the telecom implementation rather than retail. The ledger describes itself as a source ledger, not a controlled cross-provider ranking, and instructs readers to match domain, task release, agent model, user-simulator model, scaffold, prompts, trial count and pass^k before comparing rows.
| Public τ²-bench source ledger (benchlm.ai, accessed 8 August 2026) |
| GPT-5.4Published — we quoted itat the top of the public ledger | 98.9% | Different setup — not directly comparable. Measured on the telecom implementation rather than retail. The ledger describes itself as a source ledger, not a controlled cross-provider ranking, and instructs readers to match domain, task release, agent model, user-simulator model, scaffold, prompts, trial count and pass^k before comparing rows.
| Public τ²-bench source ledger (benchlm.ai, accessed 8 August 2026) |
| Claude Fable 5Published — we quoted itat the top of the public ledger | 98.5% | Different setup — not directly comparable. Measured on the telecom implementation rather than retail. The ledger describes itself as a source ledger, not a controlled cross-provider ranking, and instructs readers to match domain, task release, agent model, user-simulator model, scaffold, prompts, trial count and pass^k before comparing rows.
| Public τ²-bench source ledger (benchlm.ai, accessed 8 August 2026) |
For orientation only, repeated here so nobody has to scroll to hold both in mind: our figure on this benchmark is 61.4% over 114 scored tasks, on claude-haiku-4-5 (the cheapest and fastest model in its family). Not a ranking against the rows above.
What was thrown away
Attempted
114
Scored
114
Excluded
0
None excluded — the sample is the whole attempt.
Nothing was dropped. Every task in the domain suite was attempted and scored, including the ones that errored — a task that fails because our agent broke is a failed task, not an excluded one.
What this number can and cannot be put beside
No head-to-head is claimed here. Nothing on this page has been run through our harness with only the agent swapped, so the published figures below are context rather than a comparison — the τ² ledger varies domain, scaffold, user simulator and k between rows, and says so itself.
Comparable with
- Other runs of this same benchmark on this page, at the same modality (Text).
- Runs of this benchmark on the same arm — claude-haiku-4-5 (the cheapest and fastest model in its family). A run on a different model is a different experiment.
- Other Pass^1 figures — one attempt per task, no best-of-n.
Not comparable with
- The published figures shown for GPT-4o, Claude-3.5-Sonnet, GLM-5.2, GPT-5.4, Claude Fable 5. They are on the page for context, under a different or unstated setup — task set, scaffold, user simulator and pass^k all vary between published rows, and the public ledgers describe themselves as a record of what people reported rather than a controlled cross-provider ranking.
What passing and failing look like
No published cases for this run
The same benchmark over time
One run on record
| Run date | Score | Scored | Excluded | Ran on | Status |
|---|---|---|---|---|---|
| 31 Jul 2026 | 61.4%(70 of 114) | 114 | 0 | claude-haiku-4-5 | Published |
How this run was produced
- Agent under test
- The production Whissle agent stack — the same prompt, tool layer and guardrails a customer's agent runs on, driven over our public API.
- How it was run
- Scoring is primarily database state: did the world end up the way the task required. Concurrency 2 — higher triggers upstream rate limiting, which pollutes the score rather than measuring the agent.
- Model config
- Model claude-haiku-4-5Provider Anthropicthe cheapest and fastest model in its family, in its own vendor's lineup.The harness pinned nothing — no model, no provider, no reasoning-depth or speed setting. Those request controls did not exist on the benchmark endpoint when these runs were made, so every arm inherited the production deployment default.Read from the deployment default in force on each run date, not from the run's own record: the endpoint did not report which model served a turn until after these runs, and no run artifact names one. Automatic failover to the secondary provider (Gemini 2.5 Flash) is logged but not recorded per turn, so we cannot rule out that a small fraction of turns was served by it.
- Judge
- Independent judge — τ²-bench database-state checker. The benchmark owns the tools, the task database and the scoring. It checks whether the world ended up the way the task required — not whether a transcript reads well, and not anything we supply. We only provide the agent.
- Sampling
- Full suite — every task, no sub-selection Population: All 114 tasks in the τ²-bench retail domain. No seed — the selection was not randomised. 1 attempt per task.
- Run date
- 31 July 2026
- Harness commit
- Not recorded for this run. Newer runs pin the commit.
- Cost
- Not recorded for this run.
- Artifact
results/whissle/retail_run1.json