Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

Tau Bench

Agent · Agent or completed task · Open · Domain-general · 2024

Conversations where an agent uses domain-specific tools and must follow the relevant policies.

Reported results

pass^1 task success, retail domain · higher is betterResults 2024-10 · checked 2026-09-18
Tau Bench, retail, single run
Claude 3.5 Sonnet (Oct 2024), tool calling
69.2
Claude 3.5 Sonnet (Jun 2024), tool calling
62.6
GPT-4o, tool calling
60.4

The original, now frozen, tau-bench. Newer tau versions use different tasks and are not comparable.

Three model configurations using tools, selected from the repository’s results table. Source: tau-bench README, Sierra Research, Oct 2024.

Tau Bench, retail, four runs46.2
pass^4: success in all four of four trials, retail · higher is betterResults 2024-10 · checked 2026-09-18
Tau Bench, retail, four runs
Claude 3.5 Sonnet (Oct 2024), tool calling
46.2
Claude 3.5 Sonnet (Jun 2024), tool calling
38.7
GPT-4o, tool calling
38.3

Passing once is easier than passing repeatedly. For the October 2024 Sonnet configuration, the score falls from 69.2% on the single-run measure to 46.2% on the four-run measure.

The same three model configurations as the single-run table. Source: tau-bench README, Sierra Research, Oct 2024.

Tau Bench, airline, single run46
pass^1 task success, airline domain · higher is betterResults 2024-10 · checked 2026-09-18
Tau Bench, airline, single run
Claude 3.5 Sonnet (Oct 2024), tool calling
46
GPT-4o, tool calling
42
Claude 3.5 Sonnet (Jun 2024), tool calling
36

The airline domain has stricter policies than retail, and every system scores lower on it.

Three model configurations using tools, selected from the repository’s results table. Source: tau-bench README, Sierra Research, Oct 2024.

Tau Bench, airline, four runs22.5
pass^4: success in all four of four trials, airline · higher is betterResults 2024-10 · checked 2026-09-18
Tau Bench, airline, four runs
Claude 3.5 Sonnet (Oct 2024), tool calling
22.5
GPT-4o, tool calling
20
Claude 3.5 Sonnet (Jun 2024), tool calling
13.9

Under a quarter of airline tasks succeed four times in a row for any system shown.

The same three model configurations as the single-run table. Source: tau-bench README, Sierra Research, Oct 2024.

What the benchmark measures

Conversations where an agent uses domain-specific tools and must follow the relevant policies.

The evaluated unit is an agent or completed task. Read the source for the exact prompt, tool and harness conditions.

How it is scored

Task success rate and policy compliance score.

Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.

Sources

Tests whether agents follow rules while executing tasks.