Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

Agents

ScoredMapped, no score—Nothing recordedGapWeaknessadjacentNot an agent test, or not legalSelect any cell to open it

Agentic properties down the rows, agent benchmarks across the columns.
Legal and general agent benchmarksAdjacent evidence
Propertyfailure lives inRealm: Legalagent · US federal and stateHarvey LABagent · Primarily USLegal Research Benchagent · US federal and stateRedlineBenchagent · US commercialLegalAgentBenchagent · Chinese legalTau Benchagent · Domain-generalBFCL v4agent · Domain-generalMultiChallengemodel · Domain-generalPRBench Legalmodel · Multiple jurisdiction…
Holds state across turnsbothWeakness———————41.53Human-evaluated success—
Carries context between stagesharnessGapinside the rubricmultiple steps———————
Revises when facts or authority changebothWeakness62.1Mean weighted rubric score————————
Selects tools and builds argumentsmodel——published harness—79.1Task success rate over 300 tasks—77.5Overall accuracy——
Follows policy while actingbothGap—————46pass^1 task success———
Completes repeatedly, not onceharnessGap———3 runs, 1 for Fable—22.5pass^4: success in all four of four trials———
Completes the whole task, every criterionbothWeakness—25.42Tasks passing every rubric criterion55.29Questions passing every required rubric item50.5Turn-weighted rubric score—————
Recovers from a failed stepharnessGap—————————
Knows when to stop or hand overbothGap————————inside the rubric
Changes the system of record correctlyharnessGap—————————
Acts within authority and permissionsharnessGap—————————

Records

Every table, benchmark and gap the matrix links to.

MultiChallenge · tableMultiChallenge, inference memory41.53

Human-evaluated success · higher is better · results 2025-01 · checked 2026-09-18

Six models from 2024, general conversation.

Whether the model remembers something it worked out earlier in the conversation, rather than something it was told. Same historical cohort.

SystemScore
o1-preview41.53
Claude 3.5 Sonnet (Jun 2024)37.29
Llama 3.1 405B Instruct16.95
Gemini 1.5 Pro (Aug 2024)15.25
Mistral Large9.32
GPT-4o (Aug 2024)5.08

Compare rows within this table only. Scores are not normalised across benchmarks. Source: MultiChallenge paper, Table 2, Jan 2025. Benchmark profile.

Tau Bench · tableTau Bench, airline, single run46

pass^1 task success, airline domain · higher is better · results 2024-10 · checked 2026-09-18

Three model configurations using tools, selected from the repository’s results table.

The airline domain has stricter policies than retail, and every system scores lower on it.

SystemScore
Claude 3.5 Sonnet (Oct 2024), tool calling46
GPT-4o, tool calling42
Claude 3.5 Sonnet (Jun 2024), tool calling36

Compare rows within this table only. Scores are not normalised across benchmarks. Source: tau-bench README, Sierra Research, Oct 2024. Benchmark profile.

Tau Bench · tableTau Bench, airline, four runs22.5

pass^4: success in all four of four trials, airline · higher is better · results 2024-10 · checked 2026-09-18

The same three model configurations as the single-run table.

Under a quarter of airline tasks succeed four times in a row for any system shown.

SystemScore
Claude 3.5 Sonnet (Oct 2024), tool calling22.5
GPT-4o, tool calling20
Claude 3.5 Sonnet (Jun 2024), tool calling13.9

Compare rows within this table only. Scores are not normalised across benchmarks. Source: tau-bench README, Sierra Research, Oct 2024. Benchmark profile.

Harvey Legal Agent Benchmark · tableHarvey LAB (legal agent tasks)25.42

Tasks passing every rubric criterion · higher is better · results 2026-09 · checked 2026-09-18

Six current systems; a task passes only if every criterion passes.

Individual criterion pass rates are high, around 92–95%, but strict full-task resolution reaches only 20–25% for the leading systems. Passing most checks can still leave the task unfinished.

SystemNoteScore
Muse Spark 1.2Harvey25.42
Muse Spark 1.3 MaxHarvey23.75
Muse Spark 1.3Harvey22.08
Muse Spark 1.1Harvey20
Grok 4.6xAI15.83
Claude Fable 5Anthropic11.25

Compare rows within this table only. Scores are not normalised across benchmarks. Source: Harvey AI / VALS, Sep 2026. Benchmark profile.

RedlineBench · tableRedlineBench, overall50.5

Turn-weighted rubric score, validity gate applied · higher is better · results 2026-06 · checked 2026-09-18

Four systems, three SaaS negotiation scenarios, four turns each.

Fable 5 was tested once; the other systems were tested three times. Each of the twelve combinations of scenario and turn has equal weight in the overall score.

SystemScore
GPT-5.550.5
Claude Fable 547.3
Gemini 3.5 Flash45.1
Claude Opus 4.844.4

Compare rows within this table only. Scores are not normalised across benchmarks. Source: Crosby Intelligence, Jun 2026. Benchmark profile.

Benchmark · agent, completed task · Primarily USHarvey Legal Agent Benchmark25.42

More than 1,200 tasks spanning multiple steps across 24 practice areas, with over 75,000 assessment criteria.

Scored by

Criterion pass rate and all-criteria task completion.

Headline result

25.42 Tasks passing every rubric criterion
Muse Spark 1.2 · results 2026-09 · checked 2026-09-18
Benchmark-level context. Not a score for any row.

Supports

Performance on its long-horizon, file-based legal assignments and expert criteria under the stated harness conditions.

Does not establish

Complete coverage of legal practice or reliable operation in a particular firm’s information, supervision and production environment.

Tests sequences of legal work involving tools, research and drafting.

Primary source · first published 2026 · Full profile

Benchmark · agent, completed task · US commercialRedlineBench50.5

140 tasks involving edits to documents across three SaaS and services negotiations, each with four rounds.

Scored by

Validity gate, weighted attorney rubrics, three-judge majority vote.

Headline result

50.5 Turn-weighted rubric score, validity gate applied
GPT-5.5 · results 2026-06 · checked 2026-09-18
Benchmark-level context. Not a score for any row.

Supports

Performance on its multi-turn contract negotiations, document mechanics and attorney-authored rubrics.

Does not establish

Other agreement types, governing laws, playbooks, parties, bargaining positions or commercial contexts.

Tests DOCX redlining across negotiation rounds, including whether the edits are valid.

Primary source · first published 2025 · Full profile

Benchmark · agent, completed task · Chinese legalLegalAgentBench79.1

37 tools for interacting with legal knowledge bases.

Scored by

Task completion rate across tool-use scenarios.

Headline result

79.1 Task success rate over 300 tasks
GPT-4o with ReAct · results 2024-12
Benchmark-level context. Not a score for any row.

Supports

Completion of its Chinese legal tool-use scenarios with the available knowledge-base tools.

Does not establish

Cross-jurisdictional legal accuracy or operation with another tool and data environment.

Published at ACL 2025; tests completion of legal tasks using tools.

Primary source · first published 2025 · Full profile

Benchmark · agent, completed task · Domain-generalTau Bench69.2

Conversations where an agent uses domain-specific tools and must follow the relevant policies.

Scored by

Task success rate and policy compliance score.

Headline result

69.2 pass^1 task success, retail domain
Claude 3.5 Sonnet (Oct 2024), tool calling · results 2024-10 · checked 2026-09-18
Benchmark-level context. Not a score for any row.

Supports

Task success and policy compliance in its conversational tool-use environments.

Does not establish

Legal correctness, legal authority to act or compliance with a legal team’s specific policy.

Tests whether agents follow rules while executing tasks.

Primary source · first published 2024 · Full profile

Benchmark · agent, completed task · Domain-generalBFCL v477.5

Tool-calling tasks assessed by checking the selected functions and their arguments.

Scored by

Function correctness and argument accuracy rates.

Headline result

77.5 Overall accuracy, BFCL V4
Claude Opus 4.5, native function calling · results 2026-04
Benchmark-level context. Not a score for any row.

Supports

Function-selection and argument-construction performance on its tool-calling tasks.

Does not establish

Successful workflow completion or correctness of the legal purpose behind a call.

Standard function-calling benchmark.

Primary source · first published 2024 · Full profile

Benchmark · model or response · Domain-generalMultiChallenge

Multi-turn conversations testing whether a model keeps instructions, remembers what was inferred, stays consistent with itself and edits earlier versions reliably.

Scored by

Human-evaluated success rate per challenge category.

Headline result

No comparable table recorded.

Supports

Performance on its instruction-retention, inference-memory, consistency and version-editing challenges.

Does not establish

Legal accuracy or reliable matter-state management in production.

Tests general conversation skills relevant to negotiation and drafting. Its separate categories help examine where continuity breaks down.

Primary source · first published 2025 · Full profile

Weakness · failure lives in the model and the harnessHolds state across turnsWeakness

MultiChallenge shows models losing what they inferred and mis-editing earlier versions. That was 2024 models on hard examples. No legal agent benchmark isolates it, though Realm and RedlineBench depend on it.

Gap · failure lives in the harnessCarries context between stagesGap

Realm and Harvey LAB require it and score the whole; neither isolates whether stage two received what stage one produced. Multi-stage pipelines fail here silently: a truncated summary, a dropped attachment, a stale variable.

Weakness · failure lives in the model and the harnessRevises when facts or authority changeWeakness

Realm: Legal’s original analysis found problems selecting rules, applying facts, recognising missing information and revising later work. Newer models score higher on the mean; the failure analysis has not been repeated on them.

Gap · failure lives in the model and the harnessFollows policy while actingGap

Tau Bench (October 2024) is the only recorded test of acting under a policy: 46% airline, single run. No legal benchmark scores whether an agent kept to a playbook, a mandate or a professional rule while acting.

Gap · failure lives in the harnessCompletes repeatedly, not onceGap

Tau pass^4 roughly halves single-run scores. RedlineBench ran most systems three times and one system once, and reports a mean. No legal table in the collection reports run-to-run variance.

Weakness · failure lives in the model and the harnessCompletes the whole task, every criterionWeakness

Harvey LAB: about 94.5% of criteria pass, 25.4% of tasks complete. Legal Research Bench: 90.58% weighted credit, 55.29% strict. The gap between the two denominators is where the engineering work is.

Gap · failure lives in the harnessRecovers from a failed stepGap

Every agent benchmark here scores the end state. None records whether a system noticed a failed tool call, an empty retrieval or an exhausted budget, and what it did next. LePhantomCite’s early-termination runs are the closest observation.

Gap · failure lives in the model and the harnessKnows when to stop or hand overGap

PRBench scores uncertainty handling inside its rubric, for a model answering questions. Nothing scores an agent deciding to stop, ask or refer mid-task.

Gap · failure lives in the harnessChanges the system of record correctlyGap

Every agent benchmark here scores the answer or the document. None checks the system of record afterwards: whether the matter, the CLM or the DMS ended up in the state the task required.

Gap · failure lives in the harnessActs within authority and permissionsGap

Whether an action was within the client’s instructions, the user’s permissions and the firm’s policy is specific to you. It cannot be benchmarked publicly and it never will be.