Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

Tasks

ScoredMapped, no score—Nothing recordedGapWeaknessSelect any cell to open it

Contract review

6 steps · 5 benchmarks · 2 scored · 4 mapped only

Contract review: benchmarks across the columns, steps down the rows.
StepContractEvalmodel or responseContractNLImodel or responseContractScrubmodel or responseCUADmodel or responseRedlineBenchagent, completed task
Clause span extraction———44Precision at 80% recall—
Entailment and contradiction—Mapped———
Absence detectionGap—Mapped———
Risk identificationMapped————
Defined terms and cross-references——Mapped——
Overall redlining————50.5Turn-weighted rubric score

Records

CUAD · tableCUAD, precision at 80% recall44

Precision at 80% recall, span-level · higher is better · results 2021-03 · checked 2026-09-18

Ten fine-tuned models from the paper; original test split.

Historical baselines from 2021, not a current ranking. This measure asks how many extracted passages are correct when the model finds 80% of the relevant passages.

SystemScore
DeBERTa-xlarge44
RoBERTa-large38.1
RoBERTa-base, contracts pretraining34.1
RoBERTa-base31.1
ALBERT-xxlarge31
ALBERT-large20.9
ALBERT-xlarge20.5
ALBERT-base11.1
BERT-base8.2
BERT-large7.6

Compare rows within this table only. Scores are not normalised across benchmarks. Source: Hendrycks et al., CUAD paper, Table 2. Benchmark profile.

RedlineBench · tableRedlineBench, overall50.5

Turn-weighted rubric score, validity gate applied · higher is better · results 2026-06 · checked 2026-09-18

Four systems, three SaaS negotiation scenarios, four turns each.

Fable 5 was tested once; the other systems were tested three times. Each of the twelve combinations of scenario and turn has equal weight in the overall score.

SystemScore
GPT-5.550.5
Claude Fable 547.3
Gemini 3.5 Flash45.1
Claude Opus 4.844.4

Compare rows within this table only. Scores are not normalised across benchmarks. Source: Crosby Intelligence, Jun 2026. Benchmark profile.

LePhantomCite · tableLePhantomCite, citation verification68.8

Agentic citation-verification F1 · higher is better · results 2026-06 · checked 2026-09-18

Six agentic systems evaluated on 1,300 legal brief excerpts with injected citation problems.

F1 balances precision and recall when identifying citation problems. The paper also studies generated citations over time. Those results are separate from this verification table.

SystemNoteScore
Claude Code, Opus 4.8P 76.1 · R 62.868.8
GPT-5P 40.8 · R 84.455
Qwen3.6-27BP 28.9 · R 65.040
GPT-OSS-120BP 21.1 · R 55.130.5
Gemini 2.5 FlashP 16.9 · R 66.927
Qwen3-8BP 12.0 · R 41.118.6

Compare rows within this table only. Scores are not normalised across benchmarks. Source: LePhantomCite paper, Table 2, Jun 2026. Benchmark profile.

RedlineBench · tableRedlineBench, opening round30.3

Turn 1 rubric score · higher is better · results 2026-06 · checked 2026-09-18

Four systems, first turn only.

The first redline against a counterparty draft, before any back-and-forth. Every system scores far lower here than on its overall figure.

SystemScore
GPT-5.530.3
Claude Fable 522.6
Gemini 3.5 Flash21.9
Claude Opus 4.817.9

Compare rows within this table only. Scores are not normalised across benchmarks. Source: Crosby Intelligence, Jun 2026. Benchmark profile.

MultiChallenge · tableMultiChallenge, self-coherence45.45

Human-evaluated success · higher is better · results 2025-01 · checked 2026-09-18

Six models from 2024, general conversation.

Whether the model stays consistent with what it said earlier. Same historical cohort.

SystemScore
Claude 3.5 Sonnet (Jun 2024)45.45
o1-preview34.09
Llama 3.1 405B Instruct25
Mistral Large20.45
GPT-4o (Aug 2024)13.64
Gemini 1.5 Pro (Aug 2024)13.64

Compare rows within this table only. Scores are not normalised across benchmarks. Source: MultiChallenge paper, Table 2, Jan 2025. Benchmark profile.

MultiChallenge · tableMultiChallenge, versioned editing39.02

Human-evaluated success · higher is better · results 2025-01 · checked 2026-09-18

Six models from 2024, general conversation.

Editing an earlier version of a text correctly after later turns have changed it. The lowest category for most models in the cohort.

SystemScore
o1-preview39.02
Claude 3.5 Sonnet (Jun 2024)24.39
Gemini 1.5 Pro (Aug 2024)19.51
GPT-4o (Aug 2024)17.07
Mistral Large7.32
Llama 3.1 405B Instruct4.88

Compare rows within this table only. Scores are not normalised across benchmarks. Source: MultiChallenge paper, Table 2, Jan 2025. Benchmark profile.

MultiChallenge · tableMultiChallenge, instruction retention58.57

Human-evaluated success · higher is better · results 2025-01 · checked 2026-09-18

Six models from 2024, general conversation.

These are older models, tested on examples selected because they were difficult for them. The results help illustrate the failure mode; they do not rank current systems.

SystemScore
Claude 3.5 Sonnet (Jun 2024)58.57
o1-preview34.29
Gemini 1.5 Pro (Aug 2024)31.43
Mistral Large21.43
GPT-4o (Aug 2024)14.29
Llama 3.1 405B Instruct12.86

Compare rows within this table only. Scores are not normalised across benchmarks. Source: MultiChallenge paper, Table 2, Jan 2025. Benchmark profile.

MultiChallenge · tableMultiChallenge, inference memory41.53

Human-evaluated success · higher is better · results 2025-01 · checked 2026-09-18

Six models from 2024, general conversation.

Whether the model remembers something it worked out earlier in the conversation, rather than something it was told. Same historical cohort.

SystemScore
o1-preview41.53
Claude 3.5 Sonnet (Jun 2024)37.29
Llama 3.1 405B Instruct16.95
Gemini 1.5 Pro (Aug 2024)15.25
Mistral Large9.32
GPT-4o (Aug 2024)5.08

Compare rows within this table only. Scores are not normalised across benchmarks. Source: MultiChallenge paper, Table 2, Jan 2025. Benchmark profile.

Harvey Legal Agent Benchmark · tableHarvey LAB (legal agent tasks)25.42

Tasks passing every rubric criterion · higher is better · results 2026-09 · checked 2026-09-18

Six current systems; a task passes only if every criterion passes.

Individual criterion pass rates are high, around 92–95%, but strict full-task resolution reaches only 20–25% for the leading systems. Passing most checks can still leave the task unfinished.

SystemNoteScore
Muse Spark 1.2Harvey25.42
Muse Spark 1.3 MaxHarvey23.75
Muse Spark 1.3Harvey22.08
Muse Spark 1.1Harvey20
Grok 4.6xAI15.83
Claude Fable 5Anthropic11.25

Compare rows within this table only. Scores are not normalised across benchmarks. Source: Harvey AI / VALS, Sep 2026. Benchmark profile.

Benchmark · model or response · CommercialContractEval64.4

Clause-level risk questions derived from CUAD, tested on open and proprietary models.

Scored by

Correctness and output effectiveness scores.

Headline result

64.4 Correctness F1 on the CUAD test set
GPT-4.1 mini, zero-shot · results 2025-08
Benchmark-level context. Not a score for any row.

Supports

Performance on its clause-level risk questions and output-effectiveness criteria.

Does not establish

A complete review of an agreement or fitness for a particular organisation’s risk position.

Extends CUAD from extraction into explanation.

Primary source · first published 2025 · Full profile

Benchmark · model or response · Commercial NDAsContractNLI89.2

607 contracts tested against policy-like hypotheses, with evidence spans for each judgement.

Scored by

Three-way classification accuracy and evidence identification F1.

Headline result

89.2 NLI accuracy, macro-averaged over hypotheses
Span NLI BERT (DeBERTa-v2-xlarge backbone) · results 2021-10
Benchmark-level context. Not a score for any row.

Supports

Classification and supporting-span identification for its contract hypotheses.

Does not establish

Open-ended review, negotiation or advice against a client’s playbook.

Tests whether a contract entails, contradicts, or is silent on a policy statement.

Primary source · first published 2021 · Full profile

Benchmark · model or response · CommercialContractScrub75

Contracts prepared by lawyers with realistic errors introduced for a final review, including broken references and inconsistent defined terms.

Scored by

Macro recall by error class.

Headline result

75 Macro-average recall over nine error categories
GPT-5.5, medium reasoning · results 2026-08
Benchmark-level context. Not a score for any row.

Supports

Detection of the error classes deliberately introduced into its lawyer-prepared contracts.

Does not establish

Detection of every drafting defect or substantive legal issue in an unseen contract.

Tests the 'last pair of eyes' review. Error types drawn from real malpractice claims.

Primary source · first published 2026 · Full profile

Benchmark · model or response · US commercialCUAD44

510 contracts with 41 clause types and 13,000+ expert annotations for clause-level extraction.

Scored by

Span-level AUPR and precision at fixed recall.

Headline result

44 Precision at 80% recall, span-level
DeBERTa-xlarge · results 2021-03 · checked 2026-09-18
Benchmark-level context. Not a score for any row.

Supports

Clause identification and extraction performance across its 510 contracts and 41 clause categories.

Does not establish

Legal advice, contract risk judgment or complete contract review.

An established test of contract extraction. The results recorded here are historical baselines.

Primary source · first published 2021 · Full profile

Benchmark · agent, completed task · US commercialRedlineBench50.5

140 tasks involving edits to documents across three SaaS and services negotiations, each with four rounds.

Scored by

Validity gate, weighted attorney rubrics, three-judge majority vote.

Headline result

50.5 Turn-weighted rubric score, validity gate applied
GPT-5.5 · results 2026-06 · checked 2026-09-18
Benchmark-level context. Not a score for any row.

Supports

Performance on its multi-turn contract negotiations, document mechanics and attorney-authored rubrics.

Does not establish

Other agreement types, governing laws, playbooks, parties, bargaining positions or commercial contexts.

Tests DOCX redlining across negotiation rounds, including whether the edits are valid.

Primary source · first published 2025 · Full profile

Benchmark · model or response · English legalLegalBench-RAG

6,858 expert-annotated query-answer pairs with character-level evidence spans.

Scored by

Character-level precision and recall.

Headline result

No comparable table recorded.

Supports

Retrieval and evidence-span performance on its annotated legal query-answer pairs.

Does not establish

Complete legal research, authority validation or a supported client-facing conclusion.

Tests retrieval and grounding at character level.

Primary source · first published 2024 · Full profile

Benchmark · model or response · US litigationLePhantomCite68.8

1,300 legal brief excerpts containing deliberately introduced citation problems.

Scored by

Precision and recall by hallucination category.

Headline result

68.8 Agentic citation-verification F1
Claude Code, Opus 4.8 · results 2026-06 · checked 2026-09-18
Benchmark-level context. Not a score for any row.

Supports

Detection of the citation defects represented in its legal-brief excerpts.

Does not establish

General factual accuracy, complete authority checking or substantive correctness of the brief.

Tests whether systems can identify fabricated or otherwise problematic citations.

Primary source · first published 2026 · Full profile

Benchmark · model or response · Domain-generalRAGBench

Labelled examples for assessing retrieval-augmented generation (RAG): systems that retrieve source material before generating an answer.

Scored by

TRACe metrics for retrieval and generation quality.

Headline result

No comparable table recorded.

Supports

Retrieval and generation quality on its labelled RAG examples and TRACe measures.

Does not establish

Legal-domain accuracy or sufficiency for a particular legal research task.

Covers general domains; its evaluation approach can inform legal retrieval systems.

Primary source · first published 2024 · Full profile

Benchmark · model or response · Primarily USHELM Enterprise (Legal)

IBM extension of Stanford HELM with legal-specific scenarios.

Scored by

HELM 7-metric framework.

Headline result

No comparable table recorded.

Supports

Performance on the legal scenarios and metrics included in the HELM extension.

Does not establish

Complete legal-service quality or performance outside those scenarios.

Assesses legal scenarios using several measures of performance.

Primary source · first published 2024 · Full profile

Benchmark · model or response · Chinese legalLawBench56.3

20 tasks across three cognitive levels following Bloom's taxonomy.

Scored by

Task-specific accuracy across cognitive levels.

Headline result

56.3 Average score over 20 tasks, zero-shot
Qwen-1.5-72B-Chat · results 2024-02
Benchmark-level context. Not a score for any row.

Supports

Performance across its 20 Chinese-law tasks and three cognitive levels.

Does not establish

Cross-jurisdictional legal capability or end-to-end legal work.

Published at EMNLP 2024; covers a range of Chinese legal tasks.

Primary source · first published 2024 · Full profile

Benchmark · model or response · Primarily USLegalBench88.6

162 tasks contributed by legal professionals, covering six types of legal reasoning.

Scored by

Task-specific exact match and classification accuracy.

Headline result

88.6 Mean accuracy across 162 tasks
Claude Fable 5 · results 2026-09 · checked 2026-09-18
Benchmark-level context. Not a score for any row.

Supports

Performance on its defined mixture of legal-language and rule tasks under the recorded prompting and scoring conditions.

Does not establish

An end-to-end client matter, current legal research, completeness of a fact record, advice against client objectives, sustained matter state or safe production action.

Broad coverage of legal reasoning types. Also used in Stanford HELM.

Primary source · first published 2023 · Full profile

Benchmark · model or response · Chinese legalPLawBench69.7

Assesses language models on legal practice tasks using detailed scoring criteria.

Scored by

Rubric-guided scores across practice dimensions.

Headline result

69.7 Overall rubric score, weighted across three tasks
GPT-5.2 · results 2026-01
Benchmark-level context. Not a score for any row.

Supports

Performance on its Chinese legal-practice tasks and rubric dimensions.

Does not establish

Performance in other jurisdictions or in a particular live practice environment.

Uses assessment rubrics intended to reflect how lawyers evaluate work.

Primary source · first published 2026 · Full profile

Benchmark · model or response · Domain-generalMultiChallenge

Multi-turn conversations testing whether a model keeps instructions, remembers what was inferred, stays consistent with itself and edits earlier versions reliably.

Scored by

Human-evaluated success rate per challenge category.

Headline result

No comparable table recorded.

Supports

Performance on its instruction-retention, inference-memory, consistency and version-editing challenges.

Does not establish

Legal accuracy or reliable matter-state management in production.

Tests general conversation skills relevant to negotiation and drafting. Its separate categories help examine where continuity breaks down.

Primary source · first published 2025 · Full profile

Benchmark · agent, completed task · Domain-generalNegotiationArena

Negotiation environments covering bargaining and resource exchange between language-model agents.

Scored by

Scenario-specific negotiation outcomes.

Headline result

No comparable table recorded.

Supports

Outcomes within its bargaining and resource-exchange environments.

Does not establish

Legal negotiation quality, client alignment or performance in document-based negotiations.

Tests strategic behaviour in general negotiation settings. It does not establish whether an agent follows a legal mandate.

Primary source · first published 2024 · Full profile

Benchmark · model or response · Domain-generalLongBench v263.3

503 questions using source material ranging from 8,000 to 2 million words.

Scored by

Multiple-choice accuracy by context length.

Headline result

63.3 Overall multiple-choice accuracy, with chain of thought
Gemini 2.5 Pro · results 2025-07 · checked 2026-09-18
Benchmark-level context. Not a score for any row.

Supports

Question-answering performance over its long-context source materials.

Does not establish

Legal reasoning, source authority or dependable use of a live matter record.

Tests whether models actually use long contexts.

Primary source · first published 2024 · Full profile

Benchmark · model or response · US public M&AMAUD57.8

152 merger agreements with 92 questions and 47,000+ labels across material deal points.

Scored by

Minority-class AUPR per question.

Headline result

57.8 Mean minority-class AUPR over all questions, single-task
BigBird-base, fine-tuned · results 2023-01
Benchmark-level context. Not a score for any row.

Supports

Identification of labelled merger-agreement deal points represented in the dataset.

Does not establish

Complete M&A review, transaction strategy or advice on an unseen deal.

Tests extraction of deal points from merger agreements, making it relevant to M&A diligence.

Primary source · first published 2023 · Full profile

Benchmark · model or response · Domain-generalMP-DocVQA88.2

Questions over multi-page scanned industry documents.

Scored by

Answer exact match and page retrieval accuracy.

Headline result

88.2 ANLS on the hidden test set
RealDoc-PageTreeIndex (Zoloz) · results 2025-11
Benchmark-level context. Not a score for any row.

Supports

Answer and page-retrieval performance over its scanned multi-page documents.

Does not establish

Legal interpretation, document completeness or reliable OCR on every document type.

Relevant to reading scanned legal documents, though the benchmark itself is not law-specific.

Primary source · first published 2023 · Full profile

Benchmark · model or response · US, UK, EU, IndiaLexSumm

Eight legal summarisation datasets across multiple jurisdictions.

Scored by

Generation quality metrics requiring legal-domain supplements.

Headline result

No comparable table recorded.

Supports

Summarisation performance on its eight legal datasets under the reported metrics.

Does not establish

Accuracy for every audience, purpose, jurisdiction or downstream legal use.

Useful for comparing legal summarisation across jurisdictions.

Primary source · first published 2024 · Full profile

Benchmark · agent, completed task · Primarily USHarvey Legal Agent Benchmark25.42

More than 1,200 tasks spanning multiple steps across 24 practice areas, with over 75,000 assessment criteria.

Scored by

Criterion pass rate and all-criteria task completion.

Headline result

25.42 Tasks passing every rubric criterion
Muse Spark 1.2 · results 2026-09 · checked 2026-09-18
Benchmark-level context. Not a score for any row.

Supports

Performance on its long-horizon, file-based legal assignments and expert criteria under the stated harness conditions.

Does not establish

Complete coverage of legal practice or reliable operation in a particular firm’s information, supervision and production environment.

Tests sequences of legal work involving tools, research and drafting.

Primary source · first published 2026 · Full profile

Benchmark · agent, completed task · Domain-generalTau Bench69.2

Conversations where an agent uses domain-specific tools and must follow the relevant policies.

Scored by

Task success rate and policy compliance score.

Headline result

69.2 pass^1 task success, retail domain
Claude 3.5 Sonnet (Oct 2024), tool calling · results 2024-10 · checked 2026-09-18
Benchmark-level context. Not a score for any row.

Supports

Task success and policy compliance in its conversational tool-use environments.

Does not establish

Legal correctness, legal authority to act or compliance with a legal team’s specific policy.

Tests whether agents follow rules while executing tasks.

Primary source · first published 2024 · Full profile

Gap · partial support: ContractNLI, Realm: LegalMissing provisions and factsGap

Spotting what a document does not contain.

A review can correctly assess every clause it finds and still miss the protection that was left out. Finding what should be there needs its own test.

ContractNLI uses a NotMentioned label when a contract is silent on a fixed statement. Realm: Legal’s original diagnostic analysis also found failures when important facts were missing. Neither establishes that a system can check your full document set against a list of required provisions.

Editorial. Gap record.

Gap · no benchmark in the collectionIs the authority still good law?Gap

Whether cited authority is still good law.

A case can be real, relevant and accurately quoted, yet no longer be good law. Checking that a citation exists does not settle whether you can rely on it.

LePhantomCite checks citations and their support for a claim. Realm: Legal tests revisions when new authority or a changed timeframe is supplied. Those tests do not establish that a system can independently discover that a case has been overruled or a provision amended.

Editorial. Gap record.

Gap · partial support: Professional Reasoning Benchmark — LegalKnowing when to answer or escalateGap

Whether the system recognises uncertainty and knows when to answer, ask for missing information or hand the question to a person.

A confident wrong answer can be more dangerous than no answer. Adding “possibly” to the sentence does little to help the person relying on it.

Professional Reasoning Benchmark — Legal (PRBench) assesses how models handle uncertainty within its expert rubrics. That provides some evidence, but it does not separately test the workflow decision to answer, request information or refer the matter to a person.

Editorial. Gap record.

Gap · no benchmark in the collectionCommercial judgementGap

Whether a clause is acceptable for this deal, given its value, duration, counterparty and the client’s priorities.

The same clause can carry very different risks in a short pilot and a ten-year outsourcing deal. A playbook cannot settle every trade-off.

The benchmarks here do not directly test when commercial context justifies departing from a standard position. RedlineBench applies a playbook, which covers part of the work but leaves that judgement open.

Editorial. Gap record.

Gap · no benchmark in the collectionHolding a position under challengeGap

Whether the system changes a supported analysis just because the user pushes back.

“Are you sure?” should prompt a check of the reasoning. It should not be enough, on its own, to reverse the conclusion.

No benchmark reviewed here directly isolates pushback without new evidence. RedlineBench includes several turns, but those turns involve negotiation and counter-proposals.

Editorial. Gap record.

Gap · no benchmark in the collectionWhich document governs?Gap

Whether the system resolves the relationship between the master agreement, amendments, side letters and schedules.

The answer may be in the third amendment, not the master agreement. Reading one file can give you a clear answer to the wrong version of the deal.

The contract tests reviewed here do not establish reliable precedence across a complete agreement set. LongBench v2 covers general reasoning across documents, but it does not isolate legal amendment chains or order-of-precedence clauses.

Editorial. Gap record.

Gap · partial support: Tau BenchRun-to-run consistencyGap

Whether repeating the same task with the same documents and prompt produces materially consistent findings.

If two reviews of the same contract disagree, you need to know why. Otherwise it is hard to explain which findings someone should rely on.

Tau Bench measures success across repeated attempts at retail and airline tasks, using pass^k. That is useful background for legal systems, which need their own repeated-run tests. ContractScrub uses “consistency” differently: it checks terms within a document.

Editorial. Gap record.

Weakness · LePhantomCite, reported agent runs.Valid citations flagged as falseWeakness

Some tested agents flagged valid citations that were absent from CourtListener. A missing lookup became a false alarm.

LePhantomCite

Weakness · LePhantomCite, GPT-5 with BOED.Verification stopped with items uncheckedWeakness

GPT-5 agent misses included exhausted step budgets, repeated searches and early termination.

LePhantomCite

Weakness · Harvey LAB and Legal Research Bench, VALS AI, September 2026.Partial success does not aggregate to completionWeakness

Harvey LAB: around 94.5% of individual criteria pass, 25.4% of tasks complete. Legal Research Bench: 90.58% under weighted credit, 55.29% when every required item must pass.

Harvey Legal Agent Benchmark

Weakness · Realm: Legal, micro1.Failures recognising missing information and revising later workWeakness

Realm: Legal’s original three-model analysis found problems selecting rules, applying facts, recognising missing information and revising later work. The analysis has not been repeated on newer entries.

Realm: Legal