Tau Bench
Conversations where an agent uses domain-specific tools and must follow the relevant policies.
Reported results
| Claude 3.5 Sonnet (Oct 2024), tool calling | 69.2 | |
|---|---|---|
| Claude 3.5 Sonnet (Jun 2024), tool calling | 62.6 | |
| GPT-4o, tool calling | 60.4 |
The original, now frozen, tau-bench. Newer tau versions use different tasks and are not comparable.
Tau Bench, retail, four runs46.2
| Claude 3.5 Sonnet (Oct 2024), tool calling | 46.2 | |
|---|---|---|
| Claude 3.5 Sonnet (Jun 2024), tool calling | 38.7 | |
| GPT-4o, tool calling | 38.3 |
Passing once is easier than passing repeatedly. For the October 2024 Sonnet configuration, the score falls from 69.2% on the single-run measure to 46.2% on the four-run measure.
Tau Bench, airline, single run46
| Claude 3.5 Sonnet (Oct 2024), tool calling | 46 | |
|---|---|---|
| GPT-4o, tool calling | 42 | |
| Claude 3.5 Sonnet (Jun 2024), tool calling | 36 |
The airline domain has stricter policies than retail, and every system scores lower on it.
Tau Bench, airline, four runs22.5
| Claude 3.5 Sonnet (Oct 2024), tool calling | 22.5 | |
|---|---|---|
| GPT-4o, tool calling | 20 | |
| Claude 3.5 Sonnet (Jun 2024), tool calling | 13.9 |
Under a quarter of airline tasks succeed four times in a row for any system shown.
What the benchmark measures
Conversations where an agent uses domain-specific tools and must follow the relevant policies.
The evaluated unit is an agent or completed task. Read the source for the exact prompt, tool and harness conditions.
How it is scored
Task success rate and policy compliance score.
Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.
Sources
- Primary benchmark source
- tau-bench README, Sierra Research, Oct 2024 · results 2024-10 · checked 2026-09-18
- tau-bench README, Sierra Research, Oct 2024 · results 2024-10 · checked 2026-09-18
- tau-bench README, Sierra Research, Oct 2024 · results 2024-10 · checked 2026-09-18
- tau-bench README, Sierra Research, Oct 2024 · results 2024-10 · checked 2026-09-18
Tests whether agents follow rules while executing tasks.