Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

Skills

ScoredMapped, no score—Nothing recordedGapWeaknessSelect any cell to open it

Legal reasoning

13 skills · 15 benchmarks · 0 scored · 13 mapped only

Legal reasoning: benchmarks across the columns, skills down the rows.
SkillArtificial Analysis Legal Indexmodel or response60Occupationally weighted composite scoreCLAUSEmodel or responseno tableContractEvalmodel or response64.4Correctness F1 on the CUAD test setContractNLImodel or response89.2NLI accuracyHELM Enterprise (Legal)model or responseno tableLawBenchmodel or response56.3Average score over 20 tasksLegalBenchmodel or response88.6Mean accuracy across 162 tasksLegalLensmodel or response85.3Macro F1 on the NLI subtaskMAUDmodel or response57.8Mean minority-class AUPR over all questionsPLawBenchmodel or response69.7Overall rubric scorePRBench Legalmodel or responseno tableBFCL v4agent, completed task77.5Overall accuracyHarvey LABagent, completed task25.42Tasks passing every rubric criterionLegalAgentBenchagent, completed task79.1Task success rate over 300 tasksRealm: Legalagent, completed task62.1Mean weighted rubric score
Argument construction———————————Mapped———
Entailment———Mapped———Mapped———————
Issue spotting————Mapped—Mapped———————Mapped
Multi-step reasoning————————————MappedMappedMapped
Reading comprehension————————Mapped——————
ReasoningMapped——————————————
Rule application——————Mapped——Mapped————Mapped
Rule recall—————MappedMapped————————
Statutory interpretation—————MappedMapped————————

A tag records that the benchmark authors say their task exercises the skill. A scored cell means a comparison table is linked to that skill.

Records

LePhantomCite · tableLePhantomCite, citation verification68.8

Agentic citation-verification F1 · higher is better · results 2026-06 · checked 2026-09-18

Six agentic systems evaluated on 1,300 legal brief excerpts with injected citation problems.

F1 balances precision and recall when identifying citation problems. The paper also studies generated citations over time. Those results are separate from this verification table.

SystemNoteScore
Claude Code, Opus 4.8P 76.1 · R 62.868.8
GPT-5P 40.8 · R 84.455
Qwen3.6-27BP 28.9 · R 65.040
GPT-OSS-120BP 21.1 · R 55.130.5
Gemini 2.5 FlashP 16.9 · R 66.927
Qwen3-8BP 12.0 · R 41.118.6

Compare rows within this table only. Scores are not normalised across benchmarks. Source: LePhantomCite paper, Table 2, Jun 2026. Benchmark profile.

MultiChallenge · tableMultiChallenge, versioned editing39.02

Human-evaluated success · higher is better · results 2025-01 · checked 2026-09-18

Six models from 2024, general conversation.

Editing an earlier version of a text correctly after later turns have changed it. The lowest category for most models in the cohort.

SystemScore
o1-preview39.02
Claude 3.5 Sonnet (Jun 2024)24.39
Gemini 1.5 Pro (Aug 2024)19.51
GPT-4o (Aug 2024)17.07
Mistral Large7.32
Llama 3.1 405B Instruct4.88

Compare rows within this table only. Scores are not normalised across benchmarks. Source: MultiChallenge paper, Table 2, Jan 2025. Benchmark profile.

MultiChallenge · tableMultiChallenge, inference memory41.53

Human-evaluated success · higher is better · results 2025-01 · checked 2026-09-18

Six models from 2024, general conversation.

Whether the model remembers something it worked out earlier in the conversation, rather than something it was told. Same historical cohort.

SystemScore
o1-preview41.53
Claude 3.5 Sonnet (Jun 2024)37.29
Llama 3.1 405B Instruct16.95
Gemini 1.5 Pro (Aug 2024)15.25
Mistral Large9.32
GPT-4o (Aug 2024)5.08

Compare rows within this table only. Scores are not normalised across benchmarks. Source: MultiChallenge paper, Table 2, Jan 2025. Benchmark profile.

MultiChallenge · tableMultiChallenge, instruction retention58.57

Human-evaluated success · higher is better · results 2025-01 · checked 2026-09-18

Six models from 2024, general conversation.

These are older models, tested on examples selected because they were difficult for them. The results help illustrate the failure mode; they do not rank current systems.

SystemScore
Claude 3.5 Sonnet (Jun 2024)58.57
o1-preview34.29
Gemini 1.5 Pro (Aug 2024)31.43
Mistral Large21.43
GPT-4o (Aug 2024)14.29
Llama 3.1 405B Instruct12.86

Compare rows within this table only. Scores are not normalised across benchmarks. Source: MultiChallenge paper, Table 2, Jan 2025. Benchmark profile.

MultiChallenge · tableMultiChallenge, self-coherence45.45

Human-evaluated success · higher is better · results 2025-01 · checked 2026-09-18

Six models from 2024, general conversation.

Whether the model stays consistent with what it said earlier. Same historical cohort.

SystemScore
Claude 3.5 Sonnet (Jun 2024)45.45
o1-preview34.09
Llama 3.1 405B Instruct25
Mistral Large20.45
GPT-4o (Aug 2024)13.64
Gemini 1.5 Pro (Aug 2024)13.64

Compare rows within this table only. Scores are not normalised across benchmarks. Source: MultiChallenge paper, Table 2, Jan 2025. Benchmark profile.

Benchmark · model or response · CommercialCLAUSE

More than 7,500 contracts with deliberate changes across ten categories of anomaly.

Scored by

Detection accuracy and explanation quality.

Headline result

No comparable table recorded.

Supports

Detection and explanation of its ten categories of deliberate contractual anomaly.

Does not establish

General contract correctness or reliable legal review outside the tested anomaly set.

Tests whether models find and explain deliberately introduced errors.

Primary source · first published 2025 · Full profile

Benchmark · model or response · CommercialContractEval64.4

Clause-level risk questions derived from CUAD, tested on open and proprietary models.

Scored by

Correctness and output effectiveness scores.

Headline result

64.4 Correctness F1 on the CUAD test set
GPT-4.1 mini, zero-shot · results 2025-08
Benchmark-level context. Not a score for any row.

Supports

Performance on its clause-level risk questions and output-effectiveness criteria.

Does not establish

A complete review of an agreement or fitness for a particular organisation’s risk position.

Extends CUAD from extraction into explanation.

Primary source · first published 2025 · Full profile

Benchmark · model or response · Commercial NDAsContractNLI89.2

607 contracts tested against policy-like hypotheses, with evidence spans for each judgement.

Scored by

Three-way classification accuracy and evidence identification F1.

Headline result

89.2 NLI accuracy, macro-averaged over hypotheses
Span NLI BERT (DeBERTa-v2-xlarge backbone) · results 2021-10
Benchmark-level context. Not a score for any row.

Supports

Classification and supporting-span identification for its contract hypotheses.

Does not establish

Open-ended review, negotiation or advice against a client’s playbook.

Tests whether a contract entails, contradicts, or is silent on a policy statement.

Primary source · first published 2021 · Full profile

Benchmark · model or response · Primarily USHELM Enterprise (Legal)

IBM extension of Stanford HELM with legal-specific scenarios.

Scored by

HELM 7-metric framework.

Headline result

No comparable table recorded.

Supports

Performance on the legal scenarios and metrics included in the HELM extension.

Does not establish

Complete legal-service quality or performance outside those scenarios.

Assesses legal scenarios using several measures of performance.

Primary source · first published 2024 · Full profile

Benchmark · model or response · Chinese legalLawBench56.3

20 tasks across three cognitive levels following Bloom's taxonomy.

Scored by

Task-specific accuracy across cognitive levels.

Headline result

56.3 Average score over 20 tasks, zero-shot
Qwen-1.5-72B-Chat · results 2024-02
Benchmark-level context. Not a score for any row.

Supports

Performance across its 20 Chinese-law tasks and three cognitive levels.

Does not establish

Cross-jurisdictional legal capability or end-to-end legal work.

Published at EMNLP 2024; covers a range of Chinese legal tasks.

Primary source · first published 2024 · Full profile

Benchmark · model or response · Primarily USLegalBench88.6

162 tasks contributed by legal professionals, covering six types of legal reasoning.

Scored by

Task-specific exact match and classification accuracy.

Headline result

88.6 Mean accuracy across 162 tasks
Claude Fable 5 · results 2026-09 · checked 2026-09-18
Benchmark-level context. Not a score for any row.

Supports

Performance on its defined mixture of legal-language and rule tasks under the recorded prompting and scoring conditions.

Does not establish

An end-to-end client matter, current legal research, completeness of a fact record, advice against client objectives, sustained matter state or safe production action.

Broad coverage of legal reasoning types. Also used in Stanford HELM.

Primary source · first published 2023 · Full profile

Benchmark · model or response · US consumerLegalLens85.3

Tests identification of legal violations in text through named entity recognition (NER) and natural language inference (NLI).

Scored by

Weighted F1 for NER, macro F1 for NLI.

Headline result

85.3 Macro F1 on the NLI subtask, hidden test set
Team 1-800-Shared-Tasks (fine-tuned Phi-3 Medium) · results 2024-10
Benchmark-level context. Not a score for any row.

Supports

Violation identification performance under its named-entity and inference tasks.

Does not establish

A complete legal assessment of the underlying conduct or text.

NLLP Workshop 2024 shared task.

Primary source · first published 2024 · Full profile

Benchmark · model or response · US public M&AMAUD57.8

152 merger agreements with 92 questions and 47,000+ labels across material deal points.

Scored by

Minority-class AUPR per question.

Headline result

57.8 Mean minority-class AUPR over all questions, single-task
BigBird-base, fine-tuned · results 2023-01
Benchmark-level context. Not a score for any row.

Supports

Identification of labelled merger-agreement deal points represented in the dataset.

Does not establish

Complete M&A review, transaction strategy or advice on an unseen deal.

Tests extraction of deal points from merger agreements, making it relevant to M&A diligence.

Primary source · first published 2023 · Full profile

Benchmark · model or response · Chinese legalPLawBench69.7

Assesses language models on legal practice tasks using detailed scoring criteria.

Scored by

Rubric-guided scores across practice dimensions.

Headline result

69.7 Overall rubric score, weighted across three tasks
GPT-5.2 · results 2026-01
Benchmark-level context. Not a score for any row.

Supports

Performance on its Chinese legal-practice tasks and rubric dimensions.

Does not establish

Performance in other jurisdictions or in a particular live practice environment.

Uses assessment rubrics intended to reflect how lawyers evaluate work.

Primary source · first published 2026 · Full profile

Benchmark · agent, completed task · Domain-generalBFCL v477.5

Tool-calling tasks assessed by checking the selected functions and their arguments.

Scored by

Function correctness and argument accuracy rates.

Headline result

77.5 Overall accuracy, BFCL V4
Claude Opus 4.5, native function calling · results 2026-04
Benchmark-level context. Not a score for any row.

Supports

Function-selection and argument-construction performance on its tool-calling tasks.

Does not establish

Successful workflow completion or correctness of the legal purpose behind a call.

Standard function-calling benchmark.

Primary source · first published 2024 · Full profile

Benchmark · agent, completed task · Primarily USHarvey Legal Agent Benchmark25.42

More than 1,200 tasks spanning multiple steps across 24 practice areas, with over 75,000 assessment criteria.

Scored by

Criterion pass rate and all-criteria task completion.

Headline result

25.42 Tasks passing every rubric criterion
Muse Spark 1.2 · results 2026-09 · checked 2026-09-18
Benchmark-level context. Not a score for any row.

Supports

Performance on its long-horizon, file-based legal assignments and expert criteria under the stated harness conditions.

Does not establish

Complete coverage of legal practice or reliable operation in a particular firm’s information, supervision and production environment.

Tests sequences of legal work involving tools, research and drafting.

Primary source · first published 2026 · Full profile

Benchmark · agent, completed task · Chinese legalLegalAgentBench79.1

37 tools for interacting with legal knowledge bases.

Scored by

Task completion rate across tool-use scenarios.

Headline result

79.1 Task success rate over 300 tasks
GPT-4o with ReAct · results 2024-12
Benchmark-level context. Not a score for any row.

Supports

Completion of its Chinese legal tool-use scenarios with the available knowledge-base tools.

Does not establish

Cross-jurisdictional legal accuracy or operation with another tool and data environment.

Published at ACL 2025; tests completion of legal tasks using tools.

Primary source · first published 2025 · Full profile

Benchmark · model or response · English legalLegalBench-RAG

6,858 expert-annotated query-answer pairs with character-level evidence spans.

Scored by

Character-level precision and recall.

Headline result

No comparable table recorded.

Supports

Retrieval and evidence-span performance on its annotated legal query-answer pairs.

Does not establish

Complete legal research, authority validation or a supported client-facing conclusion.

Tests retrieval and grounding at character level.

Primary source · first published 2024 · Full profile

Benchmark · model or response · US litigationLePhantomCite68.8

1,300 legal brief excerpts containing deliberately introduced citation problems.

Scored by

Precision and recall by hallucination category.

Headline result

68.8 Agentic citation-verification F1
Claude Code, Opus 4.8 · results 2026-06 · checked 2026-09-18
Benchmark-level context. Not a score for any row.

Supports

Detection of the citation defects represented in its legal-brief excerpts.

Does not establish

General factual accuracy, complete authority checking or substantive correctness of the brief.

Tests whether systems can identify fabricated or otherwise problematic citations.

Primary source · first published 2026 · Full profile

Benchmark · model or response · US, UK, EU, IndiaLexSumm

Eight legal summarisation datasets across multiple jurisdictions.

Scored by

Generation quality metrics requiring legal-domain supplements.

Headline result

No comparable table recorded.

Supports

Summarisation performance on its eight legal datasets under the reported metrics.

Does not establish

Accuracy for every audience, purpose, jurisdiction or downstream legal use.

Useful for comparing legal summarisation across jurisdictions.

Primary source · first published 2024 · Full profile

Benchmark · model or response · Domain-generalRAGBench

Labelled examples for assessing retrieval-augmented generation (RAG): systems that retrieve source material before generating an answer.

Scored by

TRACe metrics for retrieval and generation quality.

Headline result

No comparable table recorded.

Supports

Retrieval and generation quality on its labelled RAG examples and TRACe measures.

Does not establish

Legal-domain accuracy or sufficiency for a particular legal research task.

Covers general domains; its evaluation approach can inform legal retrieval systems.

Primary source · first published 2024 · Full profile

Benchmark · model or response · CommercialContractScrub75

Contracts prepared by lawyers with realistic errors introduced for a final review, including broken references and inconsistent defined terms.

Scored by

Macro recall by error class.

Headline result

75 Macro-average recall over nine error categories
GPT-5.5, medium reasoning · results 2026-08
Benchmark-level context. Not a score for any row.

Supports

Detection of the error classes deliberately introduced into its lawyer-prepared contracts.

Does not establish

Detection of every drafting defect or substantive legal issue in an unseen contract.

Tests the 'last pair of eyes' review. Error types drawn from real malpractice claims.

Primary source · first published 2026 · Full profile

Benchmark · model or response · US commercialCUAD44

510 contracts with 41 clause types and 13,000+ expert annotations for clause-level extraction.

Scored by

Span-level AUPR and precision at fixed recall.

Headline result

44 Precision at 80% recall, span-level
DeBERTa-xlarge · results 2021-03 · checked 2026-09-18
Benchmark-level context. Not a score for any row.

Supports

Clause identification and extraction performance across its 510 contracts and 41 clause categories.

Does not establish

Legal advice, contract risk judgment or complete contract review.

An established test of contract extraction. The results recorded here are historical baselines.

Primary source · first published 2021 · Full profile

Benchmark · model or response · Domain-generalMultiChallenge

Multi-turn conversations testing whether a model keeps instructions, remembers what was inferred, stays consistent with itself and edits earlier versions reliably.

Scored by

Human-evaluated success rate per challenge category.

Headline result

No comparable table recorded.

Supports

Performance on its instruction-retention, inference-memory, consistency and version-editing challenges.

Does not establish

Legal accuracy or reliable matter-state management in production.

Tests general conversation skills relevant to negotiation and drafting. Its separate categories help examine where continuity breaks down.

Primary source · first published 2025 · Full profile

Benchmark · agent, completed task · US commercialRedlineBench50.5

140 tasks involving edits to documents across three SaaS and services negotiations, each with four rounds.

Scored by

Validity gate, weighted attorney rubrics, three-judge majority vote.

Headline result

50.5 Turn-weighted rubric score, validity gate applied
GPT-5.5 · results 2026-06 · checked 2026-09-18
Benchmark-level context. Not a score for any row.

Supports

Performance on its multi-turn contract negotiations, document mechanics and attorney-authored rubrics.

Does not establish

Other agreement types, governing laws, playbooks, parties, bargaining positions or commercial contexts.

Tests DOCX redlining across negotiation rounds, including whether the edits are valid.

Primary source · first published 2025 · Full profile

Benchmark · model or response · EU, US, internationalLexGLUE79.8

Seven datasets covering case law and legislation classification.

Scored by

Micro and macro F1 across classification tasks.

Headline result

79.8 Arithmetic mean of micro-F1 across seven tasks
Legal-BERT, fine-tuned per task · results 2022-05
Benchmark-level context. Not a score for any row.

Supports

Classification performance across its seven legal-language datasets.

Does not establish

Open-ended legal analysis, advice or current legal research.

Brings several legal language-understanding tasks into a common evaluation framework.

Primary source · first published 2022 · Full profile

Benchmark · agent, completed task · Domain-generalNegotiationArena

Negotiation environments covering bargaining and resource exchange between language-model agents.

Scored by

Scenario-specific negotiation outcomes.

Headline result

No comparable table recorded.

Supports

Outcomes within its bargaining and resource-exchange environments.

Does not establish

Legal negotiation quality, client alignment or performance in document-based negotiations.

Tests strategic behaviour in general negotiation settings. It does not establish whether an agent follows a legal mandate.

Primary source · first published 2024 · Full profile

Benchmark · model or response · ECtHR, US, Swiss, Chinese, IndianFairLex

Fairness benchmark across five legal systems testing demographic parity and equal opportunity over protected attributes.

Scored by

Group fairness metrics (demographic parity, equal opportunity) alongside task accuracy.

Headline result

No comparable table recorded.

Supports

The reported accuracy and group-fairness behaviour on its five legal-system datasets.

Does not establish

Absence of unfairness in another dataset, workflow or deployment population.

Tests classification fairness across several legal systems. It does not directly assess whether generated drafting changes with party demographics.

Primary source · first published 2022 · Full profile

Benchmark · model or response · Domain-generalLongBench v263.3

503 questions using source material ranging from 8,000 to 2 million words.

Scored by

Multiple-choice accuracy by context length.

Headline result

63.3 Overall multiple-choice accuracy, with chain of thought
Gemini 2.5 Pro · results 2025-07 · checked 2026-09-18
Benchmark-level context. Not a score for any row.

Supports

Question-answering performance over its long-context source materials.

Does not establish

Legal reasoning, source authority or dependable use of a live matter record.

Tests whether models actually use long contexts.

Primary source · first published 2024 · Full profile

Benchmark · model or response · Domain-generalMP-DocVQA88.2

Questions over multi-page scanned industry documents.

Scored by

Answer exact match and page retrieval accuracy.

Headline result

88.2 ANLS on the hidden test set
RealDoc-PageTreeIndex (Zoloz) · results 2025-11
Benchmark-level context. Not a score for any row.

Supports

Answer and page-retrieval performance over its scanned multi-page documents.

Does not establish

Legal interpretation, document completeness or reliable OCR on every document type.

Relevant to reading scanned legal documents, though the benchmark itself is not law-specific.

Primary source · first published 2023 · Full profile

Benchmark · agent, completed task · Domain-generalTau Bench69.2

Conversations where an agent uses domain-specific tools and must follow the relevant policies.

Scored by

Task success rate and policy compliance score.

Headline result

69.2 pass^1 task success, retail domain
Claude 3.5 Sonnet (Oct 2024), tool calling · results 2024-10 · checked 2026-09-18
Benchmark-level context. Not a score for any row.

Supports

Task success and policy compliance in its conversational tool-use environments.

Does not establish

Legal correctness, legal authority to act or compliance with a legal team’s specific policy.

Tests whether agents follow rules while executing tasks.

Primary source · first published 2024 · Full profile

Gap · partial support: Tau BenchRun-to-run consistencyGap

Whether repeating the same task with the same documents and prompt produces materially consistent findings.

If two reviews of the same contract disagree, you need to know why. Otherwise it is hard to explain which findings someone should rely on.

Tau Bench measures success across repeated attempts at retail and airline tasks, using pass^k. That is useful background for legal systems, which need their own repeated-run tests. ContractScrub uses “consistency” differently: it checks terms within a document.

Editorial. Gap record.

Gap · partial support: FairLexBias in reviews and scoringGap

Whether party identity, the source of a draft or its existing wording changes the assessment. This includes bias in models used to score other models.

If a scoring model favours a particular writing style or system, the ranking may tell you more about the judge than the quality of the legal work.

FairLex tests demographic parity and equal opportunity in classification. The collection does not directly test whether the same clause is judged differently because of the party or the source of the draft. Where models score other models, the reliability of that judging also needs checking.

Editorial. Gap record.

Gap · no benchmark in the collectionHolding a position under challengeGap

Whether the system changes a supported analysis just because the user pushes back.

“Are you sure?” should prompt a check of the reasoning. It should not be enough, on its own, to reverse the conclusion.

No benchmark reviewed here directly isolates pushback without new evidence. RedlineBench includes several turns, but those turns involve negotiation and counter-proposals.

Editorial. Gap record.

Gap · partial support: Professional Reasoning Benchmark — LegalKnowing when to answer or escalateGap

Whether the system recognises uncertainty and knows when to answer, ask for missing information or hand the question to a person.

A confident wrong answer can be more dangerous than no answer. Adding “possibly” to the sentence does little to help the person relying on it.

Professional Reasoning Benchmark — Legal (PRBench) assesses how models handle uncertainty within its expert rubrics. That provides some evidence, but it does not separately test the workflow decision to answer, request information or refer the matter to a person.

Editorial. Gap record.

Weakness · LePhantomCite, reported agent runs.Valid citations flagged as falseWeakness

Some tested agents flagged valid citations that were absent from CourtListener. A missing lookup became a false alarm.

LePhantomCite

Weakness · LePhantomCite, GPT-5 with BOED.Verification stopped with items uncheckedWeakness

GPT-5 agent misses included exhausted step budgets, repeated searches and early termination.

LePhantomCite