Agents
ScoredMapped, no score—Nothing recordedGapWeaknessadjacentNot an agent test, or not legalSelect any cell to open it
Records
MultiChallenge · tableMultiChallenge, inference memory41.53
Human-evaluated success · higher is better · results 2025-01 · checked 2026-09-18
Six models from 2024, general conversation.
Whether the model remembers something it worked out earlier in the conversation, rather than something it was told. Same historical cohort.
| System | Score |
|---|---|
| o1-preview | 41.53 |
| Claude 3.5 Sonnet (Jun 2024) | 37.29 |
| Llama 3.1 405B Instruct | 16.95 |
| Gemini 1.5 Pro (Aug 2024) | 15.25 |
| Mistral Large | 9.32 |
| GPT-4o (Aug 2024) | 5.08 |
Realm: Legal · tableRealm: Legal, long-horizon reasoning62.1
Mean weighted rubric score · higher is better · results 2026-09 · checked 2026-09-18
The ten highest-scoring configurations from 16 reported; litigation, transactional and compliance tasks where the record changes.
The leaderboard includes newer models than the original three-model analysis. That earlier report found problems with selecting rules, applying facts, recognising missing information and revising later work. Those findings should not be attributed to the newer models without separate testing.
| System | Note | Score |
|---|---|---|
| Claude Opus 5 (max) | Anthropic | 62.1 |
| Claude Fable 5.1 (max) | Anthropic | 60.8 |
| Claude Fable 5 | Anthropic | 55.7 |
| Kimi K3 (max) | Moonshot AI | 54.1 |
| Grok 4.6 (high) | xAI | 53.8 |
| GPT-5.6 Sol (max) | OpenAI | 51.1 |
| Muse Spark 1.3 (xhigh) | Harvey | 50.3 |
| Gemini 3.8 Flash (high) | 47.9 | |
| Grok 4.5 (high) | xAI | 43.2 |
| Muse Spark 1.1 (xhigh) | Harvey | 42.3 |
Tau Bench · tableTau Bench, airline, single run46
pass^1 task success, airline domain · higher is better · results 2024-10 · checked 2026-09-18
Three model configurations using tools, selected from the repository’s results table.
The airline domain has stricter policies than retail, and every system scores lower on it.
| System | Score |
|---|---|
| Claude 3.5 Sonnet (Oct 2024), tool calling | 46 |
| GPT-4o, tool calling | 42 |
| Claude 3.5 Sonnet (Jun 2024), tool calling | 36 |
Tau Bench · tableTau Bench, airline, four runs22.5
pass^4: success in all four of four trials, airline · higher is better · results 2024-10 · checked 2026-09-18
The same three model configurations as the single-run table.
Under a quarter of airline tasks succeed four times in a row for any system shown.
| System | Score |
|---|---|
| Claude 3.5 Sonnet (Oct 2024), tool calling | 22.5 |
| GPT-4o, tool calling | 20 |
| Claude 3.5 Sonnet (Jun 2024), tool calling | 13.9 |
Harvey Legal Agent Benchmark · tableHarvey LAB (legal agent tasks)25.42
Tasks passing every rubric criterion · higher is better · results 2026-09 · checked 2026-09-18
Six current systems; a task passes only if every criterion passes.
Individual criterion pass rates are high, around 92–95%, but strict full-task resolution reaches only 20–25% for the leading systems. Passing most checks can still leave the task unfinished.
| System | Note | Score |
|---|---|---|
| Muse Spark 1.2 | Harvey | 25.42 |
| Muse Spark 1.3 Max | Harvey | 23.75 |
| Muse Spark 1.3 | Harvey | 22.08 |
| Muse Spark 1.1 | Harvey | 20 |
| Grok 4.6 | xAI | 15.83 |
| Claude Fable 5 | Anthropic | 11.25 |
Legal Research Bench · tableLegal Research Bench, strict completion55.29
Questions passing every required rubric item · higher is better · results 2026-09 · checked 2026-09-18
Top eight of 61 reported systems; US federal and state research across eight practice areas.
The top three tie at 55.29%. Claude Opus 5 reaches 90.58% under weighted partial credit but only 55.29% when every required item must pass. Conflicting-authority questions reduce scores by 6–17 points per model.
| System | Note | Score |
|---|---|---|
| Muse Spark 1.3 Max | Harvey | 55.29 |
| Claude Opus 5 | Anthropic | 55.29 |
| Claude Fable 5.1 | Anthropic | 55.29 |
| Claude Fable 5 | Anthropic | 49.52 |
| GLM-5.3 | Zhipu AI | 49.04 |
| Grok 4.6 | xAI | 48.08 |
| GPT-5.6 Sol | OpenAI | 48.08 |
| Qwen 3.8 Max | Alibaba | 47.6 |
RedlineBench · tableRedlineBench, overall50.5
Turn-weighted rubric score, validity gate applied · higher is better · results 2026-06 · checked 2026-09-18
Four systems, three SaaS negotiation scenarios, four turns each.
Fable 5 was tested once; the other systems were tested three times. Each of the twelve combinations of scenario and turn has equal weight in the overall score.
| System | Score |
|---|---|
| GPT-5.5 | 50.5 |
| Claude Fable 5 | 47.3 |
| Gemini 3.5 Flash | 45.1 |
| Claude Opus 4.8 | 44.4 |
Benchmark · agent, completed task · US federal and stateRealm: Legal62.1
Litigation, transactional and compliance tasks completed over several steps, using legal materials and tools as the facts or authority change.
Scored by
Mean score across 35–60 weighted criteria per task, organised around issue, rule, application and conclusion (IRAC). The source also reports the best of three attempts.
Headline result
62.1 Mean weighted rubric score
Supports
Performance on changing, long-horizon legal scenarios within the supplied sandbox, tools and rubric.
Does not establish
Performance with a firm’s own matter files, legal sources, templates, reviewers, permissions or production systems.
The recorded leaderboard covers 16 model configurations. Its detailed failure analysis concerns the original three models, so the IRAC breakdown should not be applied to newer entries.
Benchmark · agent, completed task · Primarily USHarvey Legal Agent Benchmark25.42
More than 1,200 tasks spanning multiple steps across 24 practice areas, with over 75,000 assessment criteria.
Scored by
Criterion pass rate and all-criteria task completion.
Headline result
25.42 Tasks passing every rubric criterion
Supports
Performance on its long-horizon, file-based legal assignments and expert criteria under the stated harness conditions.
Does not establish
Complete coverage of legal practice or reliable operation in a particular firm’s information, supervision and production environment.
Tests sequences of legal work involving tools, research and drafting.
Benchmark · agent, completed task · US federal and stateLegal Research Bench55.29
Legal research questions across eight practice areas that require agents to find and combine sources. Answers are assessed for substance and supporting authority.
Scored by
Two measures: questions passing every required item, and weighted scores that award partial credit.
Headline result
55.29 Questions passing every required rubric item
Supports
Performance on its legal research questions using the supplied tools, sources, agent configuration and scoring method.
Does not establish
Other jurisdictions, an organisation’s research sources, completeness against a live matter record or reliability of downstream client advice.
Under the strict measure, every required item must pass. The public leaderboard is updated separately from the original open-source release.
Benchmark · agent, completed task · US commercialRedlineBench50.5
140 tasks involving edits to documents across three SaaS and services negotiations, each with four rounds.
Scored by
Validity gate, weighted attorney rubrics, three-judge majority vote.
Headline result
50.5 Turn-weighted rubric score, validity gate applied
Supports
Performance on its multi-turn contract negotiations, document mechanics and attorney-authored rubrics.
Does not establish
Other agreement types, governing laws, playbooks, parties, bargaining positions or commercial contexts.
Tests DOCX redlining across negotiation rounds, including whether the edits are valid.
Benchmark · agent, completed task · Chinese legalLegalAgentBench79.1
37 tools for interacting with legal knowledge bases.
Scored by
Task completion rate across tool-use scenarios.
Headline result
79.1 Task success rate over 300 tasks
Supports
Completion of its Chinese legal tool-use scenarios with the available knowledge-base tools.
Does not establish
Cross-jurisdictional legal accuracy or operation with another tool and data environment.
Published at ACL 2025; tests completion of legal tasks using tools.
Benchmark · agent, completed task · Domain-generalTau Bench69.2
Conversations where an agent uses domain-specific tools and must follow the relevant policies.
Scored by
Task success rate and policy compliance score.
Headline result
69.2 pass^1 task success, retail domain
Supports
Task success and policy compliance in its conversational tool-use environments.
Does not establish
Legal correctness, legal authority to act or compliance with a legal team’s specific policy.
Tests whether agents follow rules while executing tasks.
Benchmark · agent, completed task · Domain-generalBFCL v477.5
Tool-calling tasks assessed by checking the selected functions and their arguments.
Scored by
Function correctness and argument accuracy rates.
Headline result
77.5 Overall accuracy, BFCL V4
Supports
Function-selection and argument-construction performance on its tool-calling tasks.
Does not establish
Successful workflow completion or correctness of the legal purpose behind a call.
Standard function-calling benchmark.
Benchmark · model or response · Domain-generalMultiChallenge
Multi-turn conversations testing whether a model keeps instructions, remembers what was inferred, stays consistent with itself and edits earlier versions reliably.
Scored by
Human-evaluated success rate per challenge category.
Headline result
No comparable table recorded.
Supports
Performance on its instruction-retention, inference-memory, consistency and version-editing challenges.
Does not establish
Legal accuracy or reliable matter-state management in production.
Tests general conversation skills relevant to negotiation and drafting. Its separate categories help examine where continuity breaks down.
Benchmark · model or response · Multiple jurisdictions; published geographic totals cover legal and financeProfessional Reasoning Benchmark — Legal
500 legal questions developed with professionals, including a harder subset of 250. Tasks can involve several turns and use 10–30 weighted assessment criteria.
Scored by
A model judge applies weighted criteria to produce a score bounded between 0 and 1. The source also reports category analysis and confidence intervals.
Headline result
No comparable table recorded.
Supports
Performance on its professional legal questions and weighted assessment criteria.
Does not establish
Complete matter performance, current research or operation in a legal-service environment.
The displayed leaderboard does not specify whether it covers the full legal set or the harder subset. The discussion of the harder subset refers to older models. No overall score is recorded here while that distinction remains unclear.
Weakness · failure lives in the model and the harnessHolds state across turnsWeakness
MultiChallenge shows models losing what they inferred and mis-editing earlier versions. That was 2024 models on hard examples. No legal agent benchmark isolates it, though Realm and RedlineBench depend on it.
Gap · failure lives in the harnessCarries context between stagesGap
Realm and Harvey LAB require it and score the whole; neither isolates whether stage two received what stage one produced. Multi-stage pipelines fail here silently: a truncated summary, a dropped attachment, a stale variable.
Weakness · failure lives in the model and the harnessRevises when facts or authority changeWeakness
Realm: Legal’s original analysis found problems selecting rules, applying facts, recognising missing information and revising later work. Newer models score higher on the mean; the failure analysis has not been repeated on them.
Gap · failure lives in the model and the harnessFollows policy while actingGap
Tau Bench (October 2024) is the only recorded test of acting under a policy: 46% airline, single run. No legal benchmark scores whether an agent kept to a playbook, a mandate or a professional rule while acting.
Gap · failure lives in the harnessCompletes repeatedly, not onceGap
Tau pass^4 roughly halves single-run scores. RedlineBench ran most systems three times and one system once, and reports a mean. No legal table in the collection reports run-to-run variance.
Weakness · failure lives in the model and the harnessCompletes the whole task, every criterionWeakness
Harvey LAB: about 94.5% of criteria pass, 25.4% of tasks complete. Legal Research Bench: 90.58% weighted credit, 55.29% strict. The gap between the two denominators is where the engineering work is.
Gap · failure lives in the harnessRecovers from a failed stepGap
Every agent benchmark here scores the end state. None records whether a system noticed a failed tool call, an empty retrieval or an exhausted budget, and what it did next. LePhantomCite’s early-termination runs are the closest observation.
Gap · failure lives in the model and the harnessKnows when to stop or hand overGap
PRBench scores uncertainty handling inside its rubric, for a model answering questions. Nothing scores an agent deciding to stop, ask or refer mid-task.
Gap · failure lives in the harnessChanges the system of record correctlyGap
Every agent benchmark here scores the answer or the document. None checks the system of record afterwards: whether the matter, the CLM or the DMS ended up in the state the task required.
Gap · failure lives in the harnessActs within authority and permissionsGap
Whether an action was within the client’s instructions, the user’s permissions and the firm’s policy is specific to you. It cannot be benchmarked publicly and it never will be.