Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

Benchmarks

30 benchmarks in this curated collection

Headline result
LegalBench Reasoning Model Primarily US 2023 88.6 Mean accuracy across 162 tasks
Realm: Legal Agent Agent US federal and state 2026 62.1 Mean weighted rubric score
Professional Reasoning Benchmark — Legal Reasoning Model Multiple jurisdictions; published geographic totals cover legal and finance 2026 —
Artificial Analysis Legal Index Reasoning Model Cross-jurisdictional composite 2026 60 Occupationally weighted composi…
CUAD Contract Model US commercial 2021 44 Precision at 80% recall
ContractNLI Contract Model Commercial NDAs 2021 89.2 NLI accuracy
MAUD Contract Model US public M&A 2023 57.8 Mean minority-class AUPR over a…
ContractEval Contract Model Commercial 2025 64.4 Correctness F1 on the CUAD test…
ContractScrub Contract Model Commercial 2026 75 Macro-average recall over nine…
CLAUSE Assurance Model Commercial 2025 —
LegalBench-RAG Research Model English legal 2024 —
LePhantomCite Assurance Model US litigation 2026 68.8 Agentic citation-verification F1
LexSumm Documents Model US, UK, EU, India 2024 —
LongBench v2 Documents Model Domain-general 2024 63.3 Overall multiple-choice accuracy
RAGBench Assurance Model Domain-general 2024 —
MP-DocVQA Documents Model Domain-general 2023 88.2 ANLS on the hidden test set
LegalLens Reasoning Model US consumer 2024 85.3 Macro F1 on the NLI subtask
LawBench Reasoning Model Chinese legal 2024 56.3 Average score over 20 tasks
LexGLUE Reasoning Model EU, US, international 2022 79.8 Arithmetic mean of micro-F1 acr…
HELM Enterprise (Legal) Reasoning Model Primarily US 2024 —
PLawBench Reasoning Model Chinese legal 2026 69.7 Overall rubric score
FairLex Assurance Model ECtHR, US, Swiss, Chinese, Indian 2022 —
RedlineBench Contract Agent US commercial 2025 50.5 Turn-weighted rubric score
Harvey Legal Agent Benchmark Agent Agent Primarily US 2026 25.42 Tasks passing every rubric crit…
Legal Research Bench Research Agent US federal and state 2026 55.29 Questions passing every require…
LegalAgentBench Agent Agent Chinese legal 2025 79.1 Task success rate over 300 tasks
Tau Bench Agent Agent Domain-general 2024 69.2 pass^1 task success
BFCL v4 Agent Agent Domain-general 2024 77.5 Overall accuracy
MultiChallenge Assurance Model Domain-general 2025 —
NegotiationArena Agent Agent Domain-general 2024 —

Records

Benchmark · model or response · Primarily USLegalBench88.6

162 tasks contributed by legal professionals, covering six types of legal reasoning.

Scored by

Task-specific exact match and classification accuracy.

Headline result

88.6 Mean accuracy across 162 tasks
Claude Fable 5 · results 2026-09 · checked 2026-09-18
Benchmark-level context. Not a score for any row.

Supports

Performance on its defined mixture of legal-language and rule tasks under the recorded prompting and scoring conditions.

Does not establish

An end-to-end client matter, current legal research, completeness of a fact record, advice against client objectives, sustained matter state or safe production action.

Broad coverage of legal reasoning types. Also used in Stanford HELM.

Primary source · first published 2023 · Full profile

Benchmark · model or response · US commercialCUAD44

510 contracts with 41 clause types and 13,000+ expert annotations for clause-level extraction.

Scored by

Span-level AUPR and precision at fixed recall.

Headline result

44 Precision at 80% recall, span-level
DeBERTa-xlarge · results 2021-03 · checked 2026-09-18
Benchmark-level context. Not a score for any row.

Supports

Clause identification and extraction performance across its 510 contracts and 41 clause categories.

Does not establish

Legal advice, contract risk judgment or complete contract review.

An established test of contract extraction. The results recorded here are historical baselines.

Primary source · first published 2021 · Full profile

Benchmark · model or response · Commercial NDAsContractNLI89.2

607 contracts tested against policy-like hypotheses, with evidence spans for each judgement.

Scored by

Three-way classification accuracy and evidence identification F1.

Headline result

89.2 NLI accuracy, macro-averaged over hypotheses
Span NLI BERT (DeBERTa-v2-xlarge backbone) · results 2021-10
Benchmark-level context. Not a score for any row.

Supports

Classification and supporting-span identification for its contract hypotheses.

Does not establish

Open-ended review, negotiation or advice against a client’s playbook.

Tests whether a contract entails, contradicts, or is silent on a policy statement.

Primary source · first published 2021 · Full profile

Benchmark · model or response · US public M&AMAUD57.8

152 merger agreements with 92 questions and 47,000+ labels across material deal points.

Scored by

Minority-class AUPR per question.

Headline result

57.8 Mean minority-class AUPR over all questions, single-task
BigBird-base, fine-tuned · results 2023-01
Benchmark-level context. Not a score for any row.

Supports

Identification of labelled merger-agreement deal points represented in the dataset.

Does not establish

Complete M&A review, transaction strategy or advice on an unseen deal.

Tests extraction of deal points from merger agreements, making it relevant to M&A diligence.

Primary source · first published 2023 · Full profile

Benchmark · model or response · CommercialContractEval64.4

Clause-level risk questions derived from CUAD, tested on open and proprietary models.

Scored by

Correctness and output effectiveness scores.

Headline result

64.4 Correctness F1 on the CUAD test set
GPT-4.1 mini, zero-shot · results 2025-08
Benchmark-level context. Not a score for any row.

Supports

Performance on its clause-level risk questions and output-effectiveness criteria.

Does not establish

A complete review of an agreement or fitness for a particular organisation’s risk position.

Extends CUAD from extraction into explanation.

Primary source · first published 2025 · Full profile

Benchmark · model or response · CommercialContractScrub75

Contracts prepared by lawyers with realistic errors introduced for a final review, including broken references and inconsistent defined terms.

Scored by

Macro recall by error class.

Headline result

75 Macro-average recall over nine error categories
GPT-5.5, medium reasoning · results 2026-08
Benchmark-level context. Not a score for any row.

Supports

Detection of the error classes deliberately introduced into its lawyer-prepared contracts.

Does not establish

Detection of every drafting defect or substantive legal issue in an unseen contract.

Tests the 'last pair of eyes' review. Error types drawn from real malpractice claims.

Primary source · first published 2026 · Full profile

Benchmark · model or response · CommercialCLAUSE

More than 7,500 contracts with deliberate changes across ten categories of anomaly.

Scored by

Detection accuracy and explanation quality.

Headline result

No comparable table recorded.

Supports

Detection and explanation of its ten categories of deliberate contractual anomaly.

Does not establish

General contract correctness or reliable legal review outside the tested anomaly set.

Tests whether models find and explain deliberately introduced errors.

Primary source · first published 2025 · Full profile

Benchmark · model or response · English legalLegalBench-RAG

6,858 expert-annotated query-answer pairs with character-level evidence spans.

Scored by

Character-level precision and recall.

Headline result

No comparable table recorded.

Supports

Retrieval and evidence-span performance on its annotated legal query-answer pairs.

Does not establish

Complete legal research, authority validation or a supported client-facing conclusion.

Tests retrieval and grounding at character level.

Primary source · first published 2024 · Full profile

Benchmark · model or response · US litigationLePhantomCite68.8

1,300 legal brief excerpts containing deliberately introduced citation problems.

Scored by

Precision and recall by hallucination category.

Headline result

68.8 Agentic citation-verification F1
Claude Code, Opus 4.8 · results 2026-06 · checked 2026-09-18
Benchmark-level context. Not a score for any row.

Supports

Detection of the citation defects represented in its legal-brief excerpts.

Does not establish

General factual accuracy, complete authority checking or substantive correctness of the brief.

Tests whether systems can identify fabricated or otherwise problematic citations.

Primary source · first published 2026 · Full profile

Benchmark · model or response · US, UK, EU, IndiaLexSumm

Eight legal summarisation datasets across multiple jurisdictions.

Scored by

Generation quality metrics requiring legal-domain supplements.

Headline result

No comparable table recorded.

Supports

Summarisation performance on its eight legal datasets under the reported metrics.

Does not establish

Accuracy for every audience, purpose, jurisdiction or downstream legal use.

Useful for comparing legal summarisation across jurisdictions.

Primary source · first published 2024 · Full profile

Benchmark · model or response · Domain-generalLongBench v263.3

503 questions using source material ranging from 8,000 to 2 million words.

Scored by

Multiple-choice accuracy by context length.

Headline result

63.3 Overall multiple-choice accuracy, with chain of thought
Gemini 2.5 Pro · results 2025-07 · checked 2026-09-18
Benchmark-level context. Not a score for any row.

Supports

Question-answering performance over its long-context source materials.

Does not establish

Legal reasoning, source authority or dependable use of a live matter record.

Tests whether models actually use long contexts.

Primary source · first published 2024 · Full profile

Benchmark · model or response · Domain-generalRAGBench

Labelled examples for assessing retrieval-augmented generation (RAG): systems that retrieve source material before generating an answer.

Scored by

TRACe metrics for retrieval and generation quality.

Headline result

No comparable table recorded.

Supports

Retrieval and generation quality on its labelled RAG examples and TRACe measures.

Does not establish

Legal-domain accuracy or sufficiency for a particular legal research task.

Covers general domains; its evaluation approach can inform legal retrieval systems.

Primary source · first published 2024 · Full profile

Benchmark · model or response · Domain-generalMP-DocVQA88.2

Questions over multi-page scanned industry documents.

Scored by

Answer exact match and page retrieval accuracy.

Headline result

88.2 ANLS on the hidden test set
RealDoc-PageTreeIndex (Zoloz) · results 2025-11
Benchmark-level context. Not a score for any row.

Supports

Answer and page-retrieval performance over its scanned multi-page documents.

Does not establish

Legal interpretation, document completeness or reliable OCR on every document type.

Relevant to reading scanned legal documents, though the benchmark itself is not law-specific.

Primary source · first published 2023 · Full profile

Benchmark · model or response · US consumerLegalLens85.3

Tests identification of legal violations in text through named entity recognition (NER) and natural language inference (NLI).

Scored by

Weighted F1 for NER, macro F1 for NLI.

Headline result

85.3 Macro F1 on the NLI subtask, hidden test set
Team 1-800-Shared-Tasks (fine-tuned Phi-3 Medium) · results 2024-10
Benchmark-level context. Not a score for any row.

Supports

Violation identification performance under its named-entity and inference tasks.

Does not establish

A complete legal assessment of the underlying conduct or text.

NLLP Workshop 2024 shared task.

Primary source · first published 2024 · Full profile

Benchmark · model or response · Chinese legalLawBench56.3

20 tasks across three cognitive levels following Bloom's taxonomy.

Scored by

Task-specific accuracy across cognitive levels.

Headline result

56.3 Average score over 20 tasks, zero-shot
Qwen-1.5-72B-Chat · results 2024-02
Benchmark-level context. Not a score for any row.

Supports

Performance across its 20 Chinese-law tasks and three cognitive levels.

Does not establish

Cross-jurisdictional legal capability or end-to-end legal work.

Published at EMNLP 2024; covers a range of Chinese legal tasks.

Primary source · first published 2024 · Full profile

Benchmark · model or response · EU, US, internationalLexGLUE79.8

Seven datasets covering case law and legislation classification.

Scored by

Micro and macro F1 across classification tasks.

Headline result

79.8 Arithmetic mean of micro-F1 across seven tasks
Legal-BERT, fine-tuned per task · results 2022-05
Benchmark-level context. Not a score for any row.

Supports

Classification performance across its seven legal-language datasets.

Does not establish

Open-ended legal analysis, advice or current legal research.

Brings several legal language-understanding tasks into a common evaluation framework.

Primary source · first published 2022 · Full profile

Benchmark · model or response · Primarily USHELM Enterprise (Legal)

IBM extension of Stanford HELM with legal-specific scenarios.

Scored by

HELM 7-metric framework.

Headline result

No comparable table recorded.

Supports

Performance on the legal scenarios and metrics included in the HELM extension.

Does not establish

Complete legal-service quality or performance outside those scenarios.

Assesses legal scenarios using several measures of performance.

Primary source · first published 2024 · Full profile

Benchmark · model or response · Chinese legalPLawBench69.7

Assesses language models on legal practice tasks using detailed scoring criteria.

Scored by

Rubric-guided scores across practice dimensions.

Headline result

69.7 Overall rubric score, weighted across three tasks
GPT-5.2 · results 2026-01
Benchmark-level context. Not a score for any row.

Supports

Performance on its Chinese legal-practice tasks and rubric dimensions.

Does not establish

Performance in other jurisdictions or in a particular live practice environment.

Uses assessment rubrics intended to reflect how lawyers evaluate work.

Primary source · first published 2026 · Full profile

Benchmark · model or response · ECtHR, US, Swiss, Chinese, IndianFairLex

Fairness benchmark across five legal systems testing demographic parity and equal opportunity over protected attributes.

Scored by

Group fairness metrics (demographic parity, equal opportunity) alongside task accuracy.

Headline result

No comparable table recorded.

Supports

The reported accuracy and group-fairness behaviour on its five legal-system datasets.

Does not establish

Absence of unfairness in another dataset, workflow or deployment population.

Tests classification fairness across several legal systems. It does not directly assess whether generated drafting changes with party demographics.

Primary source · first published 2022 · Full profile

Benchmark · agent, completed task · US commercialRedlineBench50.5

140 tasks involving edits to documents across three SaaS and services negotiations, each with four rounds.

Scored by

Validity gate, weighted attorney rubrics, three-judge majority vote.

Headline result

50.5 Turn-weighted rubric score, validity gate applied
GPT-5.5 · results 2026-06 · checked 2026-09-18
Benchmark-level context. Not a score for any row.

Supports

Performance on its multi-turn contract negotiations, document mechanics and attorney-authored rubrics.

Does not establish

Other agreement types, governing laws, playbooks, parties, bargaining positions or commercial contexts.

Tests DOCX redlining across negotiation rounds, including whether the edits are valid.

Primary source · first published 2025 · Full profile

Benchmark · agent, completed task · Primarily USHarvey Legal Agent Benchmark25.42

More than 1,200 tasks spanning multiple steps across 24 practice areas, with over 75,000 assessment criteria.

Scored by

Criterion pass rate and all-criteria task completion.

Headline result

25.42 Tasks passing every rubric criterion
Muse Spark 1.2 · results 2026-09 · checked 2026-09-18
Benchmark-level context. Not a score for any row.

Supports

Performance on its long-horizon, file-based legal assignments and expert criteria under the stated harness conditions.

Does not establish

Complete coverage of legal practice or reliable operation in a particular firm’s information, supervision and production environment.

Tests sequences of legal work involving tools, research and drafting.

Primary source · first published 2026 · Full profile

Benchmark · agent, completed task · Chinese legalLegalAgentBench79.1

37 tools for interacting with legal knowledge bases.

Scored by

Task completion rate across tool-use scenarios.

Headline result

79.1 Task success rate over 300 tasks
GPT-4o with ReAct · results 2024-12
Benchmark-level context. Not a score for any row.

Supports

Completion of its Chinese legal tool-use scenarios with the available knowledge-base tools.

Does not establish

Cross-jurisdictional legal accuracy or operation with another tool and data environment.

Published at ACL 2025; tests completion of legal tasks using tools.

Primary source · first published 2025 · Full profile

Benchmark · agent, completed task · Domain-generalTau Bench69.2

Conversations where an agent uses domain-specific tools and must follow the relevant policies.

Scored by

Task success rate and policy compliance score.

Headline result

69.2 pass^1 task success, retail domain
Claude 3.5 Sonnet (Oct 2024), tool calling · results 2024-10 · checked 2026-09-18
Benchmark-level context. Not a score for any row.

Supports

Task success and policy compliance in its conversational tool-use environments.

Does not establish

Legal correctness, legal authority to act or compliance with a legal team’s specific policy.

Tests whether agents follow rules while executing tasks.

Primary source · first published 2024 · Full profile

Benchmark · agent, completed task · Domain-generalBFCL v477.5

Tool-calling tasks assessed by checking the selected functions and their arguments.

Scored by

Function correctness and argument accuracy rates.

Headline result

77.5 Overall accuracy, BFCL V4
Claude Opus 4.5, native function calling · results 2026-04
Benchmark-level context. Not a score for any row.

Supports

Function-selection and argument-construction performance on its tool-calling tasks.

Does not establish

Successful workflow completion or correctness of the legal purpose behind a call.

Standard function-calling benchmark.

Primary source · first published 2024 · Full profile

Benchmark · model or response · Domain-generalMultiChallenge

Multi-turn conversations testing whether a model keeps instructions, remembers what was inferred, stays consistent with itself and edits earlier versions reliably.

Scored by

Human-evaluated success rate per challenge category.

Headline result

No comparable table recorded.

Supports

Performance on its instruction-retention, inference-memory, consistency and version-editing challenges.

Does not establish

Legal accuracy or reliable matter-state management in production.

Tests general conversation skills relevant to negotiation and drafting. Its separate categories help examine where continuity breaks down.

Primary source · first published 2025 · Full profile

Benchmark · agent, completed task · Domain-generalNegotiationArena

Negotiation environments covering bargaining and resource exchange between language-model agents.

Scored by

Scenario-specific negotiation outcomes.

Headline result

No comparable table recorded.

Supports

Outcomes within its bargaining and resource-exchange environments.

Does not establish

Legal negotiation quality, client alignment or performance in document-based negotiations.

Tests strategic behaviour in general negotiation settings. It does not establish whether an agent follows a legal mandate.

Primary source · first published 2024 · Full profile