Skills
- Legal reasoning
- Research and authority
- Contract analysis
- Drafting and negotiation
- Documents and extraction
- Agentic execution
- Assurance and governance
ScoredMapped, no score—Nothing recordedGapWeaknessSelect any cell to open it
Legal reasoning
| Skill | Artificial Analysis Legal Indexmodel or response60Occupationally weighted composite score | CLAUSEmodel or responseno table | ContractEvalmodel or response64.4Correctness F1 on the CUAD test set | ContractNLImodel or response89.2NLI accuracy | HELM Enterprise (Legal)model or responseno table | LawBenchmodel or response56.3Average score over 20 tasks | LegalBenchmodel or response88.6Mean accuracy across 162 tasks | LegalLensmodel or response85.3Macro F1 on the NLI subtask | MAUDmodel or response57.8Mean minority-class AUPR over all questions | PLawBenchmodel or response69.7Overall rubric score | PRBench Legalmodel or responseno table | BFCL v4agent, completed task77.5Overall accuracy | Harvey LABagent, completed task25.42Tasks passing every rubric criterion | LegalAgentBenchagent, completed task79.1Task success rate over 300 tasks | Realm: Legalagent, completed task62.1Mean weighted rubric score |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Argument construction | — | — | — | — | — | — | — | — | — | — | — | Mapped | — | — | — |
| Entailment | — | — | — | Mapped | — | — | — | Mapped | — | — | — | — | — | — | — |
| Issue spotting | — | — | — | — | Mapped | — | Mapped | — | — | — | — | — | — | — | Mapped |
| Legal explanation | — | — | Mapped | — | — | — | — | — | — | — | — | — | — | — | — |
| Legal justification | — | Mapped | — | — | — | — | — | — | — | — | — | — | — | — | — |
| Legal knowledge | Mapped | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| Legal reasoning | — | — | — | — | Mapped | Mapped | — | — | — | Mapped | Mapped | — | — | — | Mapped |
| Multi-step reasoning | — | — | — | — | — | — | — | — | — | — | — | — | Mapped | Mapped | Mapped |
| Reading comprehension | — | — | — | — | — | — | — | — | Mapped | — | — | — | — | — | — |
| Reasoning | Mapped | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| Rule application | — | — | — | — | — | — | Mapped | — | — | Mapped | — | — | — | — | Mapped |
| Rule recall | — | — | — | — | — | Mapped | Mapped | — | — | — | — | — | — | — | — |
| Statutory interpretation | — | — | — | — | — | Mapped | Mapped | — | — | — | — | — | — | — | — |
Research and authority
| Skill | Artificial Analysis Legal Indexmodel or response60Occupationally weighted composite score | ContractNLImodel or response89.2NLI accuracy | LegalBench-RAGmodel or responseno table | LePhantomCitemodel or response68.8Agentic citation-verification F1 | LexSummmodel or responseno table | RAGBenchmodel or responseno table | Legal Research Benchagent, completed task55.29Questions passing every required rubric item | LegalAgentBenchagent, completed task79.1Task success rate over 300 tasks |
|---|---|---|---|---|---|---|---|---|
| Authority selection | — | — | — | — | — | — | Mapped | — |
| Citation | — | — | — | — | — | — | Mapped | — |
| Citation verificationWeakness | — | — | — | 68.8Agentic citation-verification F1 | — | — | — | — |
| Evidence extraction | — | Mapped | — | — | — | — | — | — |
| Evidence spans | — | — | Mapped | — | — | — | — | — |
| Faithfulness | — | — | — | — | Mapped | Mapped | — | — |
| Grounding | — | — | Mapped | — | — | — | — | — |
| Legal research | — | — | — | — | — | — | Mapped | Mapped |
| Long-context evidence | Mapped | — | — | — | — | — | — | — |
| Quotation checking | — | — | — | Mapped | — | — | — | — |
| RAG evaluation | — | — | Mapped | — | — | Mapped | — | — |
| Relevance | — | — | — | — | — | Mapped | — | — |
| Retrieval | — | — | Mapped | — | — | — | — | — |
| Synthesis | — | — | — | — | — | — | Mapped | — |
Contract analysis
| Skill | ContractEvalmodel or response64.4Correctness F1 on the CUAD test set | ContractScrubmodel or response75Macro-average recall over nine error categories | CUADmodel or response44Precision at 80% recall | LegalLensmodel or response85.3Macro F1 on the NLI subtask | MAUDmodel or response57.8Mean minority-class AUPR over all questions | MultiChallengemodel or responseno table | RedlineBenchagent, completed task50.5Turn-weighted rubric score |
|---|---|---|---|---|---|---|---|
| Clause analysis | Mapped | — | — | — | — | — | — |
| Clause classification | — | — | Mapped | — | — | — | — |
| Clause extraction | — | — | Mapped | — | — | — | — |
| Contract scrubbing | — | Mapped | — | — | — | — | — |
| Cross-references | — | Mapped | — | — | — | — | — |
| Deal point extraction | — | — | — | — | Mapped | — | — |
| Defined terms | — | Mapped | — | — | — | — | — |
| Playbook application | — | — | — | — | — | — | Mapped |
| Redlining | — | — | — | — | — | — | Mapped |
| Risk identification | Mapped | — | — | — | — | — | — |
| Versioned editing | — | — | — | — | — | 39.02Human-evaluated success | — |
| Violation detection | — | — | — | Mapped | — | — | — |
Drafting and negotiation
| Skill | Artificial Analysis Legal Indexmodel or response60Occupationally weighted composite score | HELM Enterprise (Legal)model or responseno table | LexGLUEmodel or response79.8Arithmetic mean of micro-F1 across seven tasks | LexSummmodel or responseno table | PLawBenchmodel or response69.7Overall rubric score | Harvey LABagent, completed task25.42Tasks passing every rubric criterion | NegotiationArenaagent, completed taskno table | RedlineBenchagent, completed task50.5Turn-weighted rubric score |
|---|---|---|---|---|---|---|---|---|
| Audience adaptation | — | — | — | Mapped | — | — | — | — |
| Bargaining | — | — | — | — | — | — | Mapped | — |
| Drafting | — | — | — | — | Mapped | Mapped | — | — |
| Legal topic coding | — | — | Mapped | — | — | — | — | — |
| Negotiation | — | — | — | — | — | — | Mapped | Mapped |
| Professional legal tasks | Mapped | — | — | — | — | — | — | — |
| Strategic interaction | — | — | — | — | — | — | Mapped | — |
| Summarisation | — | Mapped | — | Mapped | — | — | — | — |
Documents and extraction
| Skill | CUADmodel or response44Precision at 80% recall | FairLexmodel or responseno table | HELM Enterprise (Legal)model or responseno table | LawBenchmodel or response56.3Average score over 20 tasks | LegalLensmodel or response85.3Macro F1 on the NLI subtask | LexGLUEmodel or response79.8Arithmetic mean of micro-F1 across seven tasks | LongBench v2model or response63.3Overall multiple-choice accuracy | MAUDmodel or response57.8Mean minority-class AUPR over all questions | MP-DocVQAmodel or response88.2ANLS on the hidden test set | PLawBenchmodel or response69.7Overall rubric score |
|---|---|---|---|---|---|---|---|---|---|---|
| Classification | — | Mapped | Mapped | Mapped | — | Mapped | — | Mapped | — | Mapped |
| Long context | — | — | — | — | — | — | Mapped | — | — | — |
| Multi-document reasoning | — | — | — | — | — | — | Mapped | — | — | — |
| Multi-label prediction | — | — | — | — | — | Mapped | — | — | — | — |
| Multi-page retrieval | — | — | — | — | — | — | — | — | Mapped | — |
| Named entity recognition | — | — | — | — | Mapped | — | — | — | — | — |
| OCR robustness | — | — | — | — | — | — | — | — | Mapped | — |
| Span extraction | Mapped | — | — | — | — | — | — | — | — | — |
| Visual document understanding | — | — | — | — | — | — | — | — | Mapped | — |
Agentic execution
| Skill | MultiChallengemodel or responseno table | PRBench Legalmodel or responseno table | BFCL v4agent, completed task77.5Overall accuracy | Harvey LABagent, completed task25.42Tasks passing every rubric criterion | LegalAgentBenchagent, completed task79.1Task success rate over 300 tasks | Realm: Legalagent, completed task62.1Mean weighted rubric score | Tau Benchagent, completed task69.2pass^1 task success |
|---|---|---|---|---|---|---|---|
| Agentic execution | — | — | — | Mapped | Mapped | — | Mapped |
| Function selection | — | — | Mapped | — | — | — | — |
| Inference memory | 41.53Human-evaluated success | — | — | — | — | — | — |
| Instruction retention | 58.57Human-evaluated success | — | — | — | — | — | — |
| Multi-turn reasoning | — | Mapped | — | — | — | — | — |
| Policy compliance | — | — | — | — | — | — | Mapped |
| Tool use | — | — | Mapped | Mapped | Mapped | Mapped | Mapped |
Assurance and governance
| Skill | Artificial Analysis Legal Indexmodel or response60Occupationally weighted composite score | CLAUSEmodel or responseno table | ContractNLImodel or response89.2NLI accuracy | ContractScrubmodel or response75Macro-average recall over nine error categories | FairLexmodel or responseno table | LePhantomCitemodel or response68.8Agentic citation-verification F1 | MultiChallengemodel or responseno table | PRBench Legalmodel or responseno table | RAGBenchmodel or responseno table |
|---|---|---|---|---|---|---|---|---|---|
| Adversarial testing | — | Mapped | — | — | — | — | — | — | — |
| Anomaly detection | — | Mapped | — | — | — | — | — | — | — |
| Auditability | — | — | — | — | — | — | — | Mapped | — |
| Completeness | — | — | — | — | — | — | — | — | Mapped |
| Consistency checkingGap | — | — | — | Mapped | — | — | — | — | — |
| Contradiction detection | — | — | Mapped | — | — | — | — | — | — |
| Demographic parityGap | — | — | — | — | Mapped | — | — | — | — |
| FairnessGap | — | — | — | — | Mapped | — | — | — | — |
| Hallucination detection | — | — | — | — | — | Mapped | — | — | — |
| Non-hallucination | Mapped | — | — | — | — | — | — | — | — |
| Self-coherenceGap | — | — | — | — | — | — | 45.45Human-evaluated success | — | — |
| Uncertainty handlingGap | — | — | — | — | — | — | — | Mapped | — |
Records
LePhantomCite · tableLePhantomCite, citation verification68.8
Agentic citation-verification F1 · higher is better · results 2026-06 · checked 2026-09-18
Six agentic systems evaluated on 1,300 legal brief excerpts with injected citation problems.
F1 balances precision and recall when identifying citation problems. The paper also studies generated citations over time. Those results are separate from this verification table.
| System | Note | Score |
|---|---|---|
| Claude Code, Opus 4.8 | P 76.1 · R 62.8 | 68.8 |
| GPT-5 | P 40.8 · R 84.4 | 55 |
| Qwen3.6-27B | P 28.9 · R 65.0 | 40 |
| GPT-OSS-120B | P 21.1 · R 55.1 | 30.5 |
| Gemini 2.5 Flash | P 16.9 · R 66.9 | 27 |
| Qwen3-8B | P 12.0 · R 41.1 | 18.6 |
MultiChallenge · tableMultiChallenge, versioned editing39.02
Human-evaluated success · higher is better · results 2025-01 · checked 2026-09-18
Six models from 2024, general conversation.
Editing an earlier version of a text correctly after later turns have changed it. The lowest category for most models in the cohort.
| System | Score |
|---|---|
| o1-preview | 39.02 |
| Claude 3.5 Sonnet (Jun 2024) | 24.39 |
| Gemini 1.5 Pro (Aug 2024) | 19.51 |
| GPT-4o (Aug 2024) | 17.07 |
| Mistral Large | 7.32 |
| Llama 3.1 405B Instruct | 4.88 |
MultiChallenge · tableMultiChallenge, inference memory41.53
Human-evaluated success · higher is better · results 2025-01 · checked 2026-09-18
Six models from 2024, general conversation.
Whether the model remembers something it worked out earlier in the conversation, rather than something it was told. Same historical cohort.
| System | Score |
|---|---|
| o1-preview | 41.53 |
| Claude 3.5 Sonnet (Jun 2024) | 37.29 |
| Llama 3.1 405B Instruct | 16.95 |
| Gemini 1.5 Pro (Aug 2024) | 15.25 |
| Mistral Large | 9.32 |
| GPT-4o (Aug 2024) | 5.08 |
MultiChallenge · tableMultiChallenge, instruction retention58.57
Human-evaluated success · higher is better · results 2025-01 · checked 2026-09-18
Six models from 2024, general conversation.
These are older models, tested on examples selected because they were difficult for them. The results help illustrate the failure mode; they do not rank current systems.
| System | Score |
|---|---|
| Claude 3.5 Sonnet (Jun 2024) | 58.57 |
| o1-preview | 34.29 |
| Gemini 1.5 Pro (Aug 2024) | 31.43 |
| Mistral Large | 21.43 |
| GPT-4o (Aug 2024) | 14.29 |
| Llama 3.1 405B Instruct | 12.86 |
MultiChallenge · tableMultiChallenge, self-coherence45.45
Human-evaluated success · higher is better · results 2025-01 · checked 2026-09-18
Six models from 2024, general conversation.
Whether the model stays consistent with what it said earlier. Same historical cohort.
| System | Score |
|---|---|
| Claude 3.5 Sonnet (Jun 2024) | 45.45 |
| o1-preview | 34.09 |
| Llama 3.1 405B Instruct | 25 |
| Mistral Large | 20.45 |
| GPT-4o (Aug 2024) | 13.64 |
| Gemini 1.5 Pro (Aug 2024) | 13.64 |
Benchmark · model or response · Cross-jurisdictional compositeArtificial Analysis Legal Index60
A combined score from seven evaluations, weighted to reflect the skills used in legal work.
Scored by
Weighted composite: legal knowledge 35%, agentic knowledge work 25%, reasoning 15%, long context 10%, non-hallucination 10%, tool use 5%.
Headline result
60 Occupationally weighted composite score
Supports
Performance under the index’s published combination and weighting of seven evaluations.
Does not establish
A universal legal-capability score or equivalent performance across every legal workflow.
Useful for a broad comparison, but several components test general abilities rather than legal work. The chart and FAQ name different leading configurations; the table here follows the chart.
Benchmark · model or response · CommercialCLAUSE
More than 7,500 contracts with deliberate changes across ten categories of anomaly.
Scored by
Detection accuracy and explanation quality.
Headline result
No comparable table recorded.
Supports
Detection and explanation of its ten categories of deliberate contractual anomaly.
Does not establish
General contract correctness or reliable legal review outside the tested anomaly set.
Tests whether models find and explain deliberately introduced errors.
Benchmark · model or response · CommercialContractEval64.4
Clause-level risk questions derived from CUAD, tested on open and proprietary models.
Scored by
Correctness and output effectiveness scores.
Headline result
64.4 Correctness F1 on the CUAD test set
Supports
Performance on its clause-level risk questions and output-effectiveness criteria.
Does not establish
A complete review of an agreement or fitness for a particular organisation’s risk position.
Extends CUAD from extraction into explanation.
Benchmark · model or response · Commercial NDAsContractNLI89.2
607 contracts tested against policy-like hypotheses, with evidence spans for each judgement.
Scored by
Three-way classification accuracy and evidence identification F1.
Headline result
89.2 NLI accuracy, macro-averaged over hypotheses
Supports
Classification and supporting-span identification for its contract hypotheses.
Does not establish
Open-ended review, negotiation or advice against a client’s playbook.
Tests whether a contract entails, contradicts, or is silent on a policy statement.
Benchmark · model or response · Primarily USHELM Enterprise (Legal)
IBM extension of Stanford HELM with legal-specific scenarios.
Scored by
HELM 7-metric framework.
Headline result
No comparable table recorded.
Supports
Performance on the legal scenarios and metrics included in the HELM extension.
Does not establish
Complete legal-service quality or performance outside those scenarios.
Assesses legal scenarios using several measures of performance.
Benchmark · model or response · Chinese legalLawBench56.3
20 tasks across three cognitive levels following Bloom's taxonomy.
Scored by
Task-specific accuracy across cognitive levels.
Headline result
56.3 Average score over 20 tasks, zero-shot
Supports
Performance across its 20 Chinese-law tasks and three cognitive levels.
Does not establish
Cross-jurisdictional legal capability or end-to-end legal work.
Published at EMNLP 2024; covers a range of Chinese legal tasks.
Benchmark · model or response · Primarily USLegalBench88.6
162 tasks contributed by legal professionals, covering six types of legal reasoning.
Scored by
Task-specific exact match and classification accuracy.
Headline result
88.6 Mean accuracy across 162 tasks
Supports
Performance on its defined mixture of legal-language and rule tasks under the recorded prompting and scoring conditions.
Does not establish
An end-to-end client matter, current legal research, completeness of a fact record, advice against client objectives, sustained matter state or safe production action.
Broad coverage of legal reasoning types. Also used in Stanford HELM.
Benchmark · model or response · US consumerLegalLens85.3
Tests identification of legal violations in text through named entity recognition (NER) and natural language inference (NLI).
Scored by
Weighted F1 for NER, macro F1 for NLI.
Headline result
85.3 Macro F1 on the NLI subtask, hidden test set
Supports
Violation identification performance under its named-entity and inference tasks.
Does not establish
A complete legal assessment of the underlying conduct or text.
NLLP Workshop 2024 shared task.
Benchmark · model or response · US public M&AMAUD57.8
152 merger agreements with 92 questions and 47,000+ labels across material deal points.
Scored by
Minority-class AUPR per question.
Headline result
57.8 Mean minority-class AUPR over all questions, single-task
Supports
Identification of labelled merger-agreement deal points represented in the dataset.
Does not establish
Complete M&A review, transaction strategy or advice on an unseen deal.
Tests extraction of deal points from merger agreements, making it relevant to M&A diligence.
Benchmark · model or response · Chinese legalPLawBench69.7
Assesses language models on legal practice tasks using detailed scoring criteria.
Scored by
Rubric-guided scores across practice dimensions.
Headline result
69.7 Overall rubric score, weighted across three tasks
Supports
Performance on its Chinese legal-practice tasks and rubric dimensions.
Does not establish
Performance in other jurisdictions or in a particular live practice environment.
Uses assessment rubrics intended to reflect how lawyers evaluate work.
Benchmark · model or response · Multiple jurisdictions; published geographic totals cover legal and financeProfessional Reasoning Benchmark — Legal
500 legal questions developed with professionals, including a harder subset of 250. Tasks can involve several turns and use 10–30 weighted assessment criteria.
Scored by
A model judge applies weighted criteria to produce a score bounded between 0 and 1. The source also reports category analysis and confidence intervals.
Headline result
No comparable table recorded.
Supports
Performance on its professional legal questions and weighted assessment criteria.
Does not establish
Complete matter performance, current research or operation in a legal-service environment.
The displayed leaderboard does not specify whether it covers the full legal set or the harder subset. The discussion of the harder subset refers to older models. No overall score is recorded here while that distinction remains unclear.
Benchmark · agent, completed task · Domain-generalBFCL v477.5
Tool-calling tasks assessed by checking the selected functions and their arguments.
Scored by
Function correctness and argument accuracy rates.
Headline result
77.5 Overall accuracy, BFCL V4
Supports
Function-selection and argument-construction performance on its tool-calling tasks.
Does not establish
Successful workflow completion or correctness of the legal purpose behind a call.
Standard function-calling benchmark.
Benchmark · agent, completed task · Primarily USHarvey Legal Agent Benchmark25.42
More than 1,200 tasks spanning multiple steps across 24 practice areas, with over 75,000 assessment criteria.
Scored by
Criterion pass rate and all-criteria task completion.
Headline result
25.42 Tasks passing every rubric criterion
Supports
Performance on its long-horizon, file-based legal assignments and expert criteria under the stated harness conditions.
Does not establish
Complete coverage of legal practice or reliable operation in a particular firm’s information, supervision and production environment.
Tests sequences of legal work involving tools, research and drafting.
Benchmark · agent, completed task · Chinese legalLegalAgentBench79.1
37 tools for interacting with legal knowledge bases.
Scored by
Task completion rate across tool-use scenarios.
Headline result
79.1 Task success rate over 300 tasks
Supports
Completion of its Chinese legal tool-use scenarios with the available knowledge-base tools.
Does not establish
Cross-jurisdictional legal accuracy or operation with another tool and data environment.
Published at ACL 2025; tests completion of legal tasks using tools.
Benchmark · agent, completed task · US federal and stateRealm: Legal62.1
Litigation, transactional and compliance tasks completed over several steps, using legal materials and tools as the facts or authority change.
Scored by
Mean score across 35–60 weighted criteria per task, organised around issue, rule, application and conclusion (IRAC). The source also reports the best of three attempts.
Headline result
62.1 Mean weighted rubric score
Supports
Performance on changing, long-horizon legal scenarios within the supplied sandbox, tools and rubric.
Does not establish
Performance with a firm’s own matter files, legal sources, templates, reviewers, permissions or production systems.
The recorded leaderboard covers 16 model configurations. Its detailed failure analysis concerns the original three models, so the IRAC breakdown should not be applied to newer entries.
Benchmark · model or response · English legalLegalBench-RAG
6,858 expert-annotated query-answer pairs with character-level evidence spans.
Scored by
Character-level precision and recall.
Headline result
No comparable table recorded.
Supports
Retrieval and evidence-span performance on its annotated legal query-answer pairs.
Does not establish
Complete legal research, authority validation or a supported client-facing conclusion.
Tests retrieval and grounding at character level.
Benchmark · model or response · US litigationLePhantomCite68.8
1,300 legal brief excerpts containing deliberately introduced citation problems.
Scored by
Precision and recall by hallucination category.
Headline result
68.8 Agentic citation-verification F1
Supports
Detection of the citation defects represented in its legal-brief excerpts.
Does not establish
General factual accuracy, complete authority checking or substantive correctness of the brief.
Tests whether systems can identify fabricated or otherwise problematic citations.
Benchmark · model or response · US, UK, EU, IndiaLexSumm
Eight legal summarisation datasets across multiple jurisdictions.
Scored by
Generation quality metrics requiring legal-domain supplements.
Headline result
No comparable table recorded.
Supports
Summarisation performance on its eight legal datasets under the reported metrics.
Does not establish
Accuracy for every audience, purpose, jurisdiction or downstream legal use.
Useful for comparing legal summarisation across jurisdictions.
Benchmark · model or response · Domain-generalRAGBench
Labelled examples for assessing retrieval-augmented generation (RAG): systems that retrieve source material before generating an answer.
Scored by
TRACe metrics for retrieval and generation quality.
Headline result
No comparable table recorded.
Supports
Retrieval and generation quality on its labelled RAG examples and TRACe measures.
Does not establish
Legal-domain accuracy or sufficiency for a particular legal research task.
Covers general domains; its evaluation approach can inform legal retrieval systems.
Benchmark · agent, completed task · US federal and stateLegal Research Bench55.29
Legal research questions across eight practice areas that require agents to find and combine sources. Answers are assessed for substance and supporting authority.
Scored by
Two measures: questions passing every required item, and weighted scores that award partial credit.
Headline result
55.29 Questions passing every required rubric item
Supports
Performance on its legal research questions using the supplied tools, sources, agent configuration and scoring method.
Does not establish
Other jurisdictions, an organisation’s research sources, completeness against a live matter record or reliability of downstream client advice.
Under the strict measure, every required item must pass. The public leaderboard is updated separately from the original open-source release.
Benchmark · model or response · CommercialContractScrub75
Contracts prepared by lawyers with realistic errors introduced for a final review, including broken references and inconsistent defined terms.
Scored by
Macro recall by error class.
Headline result
75 Macro-average recall over nine error categories
Supports
Detection of the error classes deliberately introduced into its lawyer-prepared contracts.
Does not establish
Detection of every drafting defect or substantive legal issue in an unseen contract.
Tests the 'last pair of eyes' review. Error types drawn from real malpractice claims.
Benchmark · model or response · US commercialCUAD44
510 contracts with 41 clause types and 13,000+ expert annotations for clause-level extraction.
Scored by
Span-level AUPR and precision at fixed recall.
Headline result
44 Precision at 80% recall, span-level
Supports
Clause identification and extraction performance across its 510 contracts and 41 clause categories.
Does not establish
Legal advice, contract risk judgment or complete contract review.
An established test of contract extraction. The results recorded here are historical baselines.
Benchmark · model or response · Domain-generalMultiChallenge
Multi-turn conversations testing whether a model keeps instructions, remembers what was inferred, stays consistent with itself and edits earlier versions reliably.
Scored by
Human-evaluated success rate per challenge category.
Headline result
No comparable table recorded.
Supports
Performance on its instruction-retention, inference-memory, consistency and version-editing challenges.
Does not establish
Legal accuracy or reliable matter-state management in production.
Tests general conversation skills relevant to negotiation and drafting. Its separate categories help examine where continuity breaks down.
Benchmark · agent, completed task · US commercialRedlineBench50.5
140 tasks involving edits to documents across three SaaS and services negotiations, each with four rounds.
Scored by
Validity gate, weighted attorney rubrics, three-judge majority vote.
Headline result
50.5 Turn-weighted rubric score, validity gate applied
Supports
Performance on its multi-turn contract negotiations, document mechanics and attorney-authored rubrics.
Does not establish
Other agreement types, governing laws, playbooks, parties, bargaining positions or commercial contexts.
Tests DOCX redlining across negotiation rounds, including whether the edits are valid.
Benchmark · model or response · EU, US, internationalLexGLUE79.8
Seven datasets covering case law and legislation classification.
Scored by
Micro and macro F1 across classification tasks.
Headline result
79.8 Arithmetic mean of micro-F1 across seven tasks
Supports
Classification performance across its seven legal-language datasets.
Does not establish
Open-ended legal analysis, advice or current legal research.
Brings several legal language-understanding tasks into a common evaluation framework.
Benchmark · agent, completed task · Domain-generalNegotiationArena
Negotiation environments covering bargaining and resource exchange between language-model agents.
Scored by
Scenario-specific negotiation outcomes.
Headline result
No comparable table recorded.
Supports
Outcomes within its bargaining and resource-exchange environments.
Does not establish
Legal negotiation quality, client alignment or performance in document-based negotiations.
Tests strategic behaviour in general negotiation settings. It does not establish whether an agent follows a legal mandate.
Benchmark · model or response · ECtHR, US, Swiss, Chinese, IndianFairLex
Fairness benchmark across five legal systems testing demographic parity and equal opportunity over protected attributes.
Scored by
Group fairness metrics (demographic parity, equal opportunity) alongside task accuracy.
Headline result
No comparable table recorded.
Supports
The reported accuracy and group-fairness behaviour on its five legal-system datasets.
Does not establish
Absence of unfairness in another dataset, workflow or deployment population.
Tests classification fairness across several legal systems. It does not directly assess whether generated drafting changes with party demographics.
Benchmark · model or response · Domain-generalLongBench v263.3
503 questions using source material ranging from 8,000 to 2 million words.
Scored by
Multiple-choice accuracy by context length.
Headline result
63.3 Overall multiple-choice accuracy, with chain of thought
Supports
Question-answering performance over its long-context source materials.
Does not establish
Legal reasoning, source authority or dependable use of a live matter record.
Tests whether models actually use long contexts.
Benchmark · model or response · Domain-generalMP-DocVQA88.2
Questions over multi-page scanned industry documents.
Scored by
Answer exact match and page retrieval accuracy.
Headline result
88.2 ANLS on the hidden test set
Supports
Answer and page-retrieval performance over its scanned multi-page documents.
Does not establish
Legal interpretation, document completeness or reliable OCR on every document type.
Relevant to reading scanned legal documents, though the benchmark itself is not law-specific.
Benchmark · agent, completed task · Domain-generalTau Bench69.2
Conversations where an agent uses domain-specific tools and must follow the relevant policies.
Scored by
Task success rate and policy compliance score.
Headline result
69.2 pass^1 task success, retail domain
Supports
Task success and policy compliance in its conversational tool-use environments.
Does not establish
Legal correctness, legal authority to act or compliance with a legal team’s specific policy.
Tests whether agents follow rules while executing tasks.
Gap · partial support: Tau BenchRun-to-run consistencyGap
Whether repeating the same task with the same documents and prompt produces materially consistent findings.
If two reviews of the same contract disagree, you need to know why. Otherwise it is hard to explain which findings someone should rely on.
Tau Bench measures success across repeated attempts at retail and airline tasks, using pass^k. That is useful background for legal systems, which need their own repeated-run tests. ContractScrub uses “consistency” differently: it checks terms within a document.
Gap · partial support: FairLexBias in reviews and scoringGap
Whether party identity, the source of a draft or its existing wording changes the assessment. This includes bias in models used to score other models.
If a scoring model favours a particular writing style or system, the ranking may tell you more about the judge than the quality of the legal work.
FairLex tests demographic parity and equal opportunity in classification. The collection does not directly test whether the same clause is judged differently because of the party or the source of the draft. Where models score other models, the reliability of that judging also needs checking.
Gap · no benchmark in the collectionHolding a position under challengeGap
Whether the system changes a supported analysis just because the user pushes back.
“Are you sure?” should prompt a check of the reasoning. It should not be enough, on its own, to reverse the conclusion.
No benchmark reviewed here directly isolates pushback without new evidence. RedlineBench includes several turns, but those turns involve negotiation and counter-proposals.
Gap · partial support: Professional Reasoning Benchmark — LegalKnowing when to answer or escalateGap
Whether the system recognises uncertainty and knows when to answer, ask for missing information or hand the question to a person.
A confident wrong answer can be more dangerous than no answer. Adding “possibly” to the sentence does little to help the person relying on it.
Professional Reasoning Benchmark — Legal (PRBench) assesses how models handle uncertainty within its expert rubrics. That provides some evidence, but it does not separately test the workflow decision to answer, request information or refer the matter to a person.
Weakness · LePhantomCite, reported agent runs.Valid citations flagged as falseWeakness
Some tested agents flagged valid citations that were absent from CourtListener. A missing lookup became a false alarm.
Weakness · LePhantomCite, GPT-5 with BOED.Verification stopped with items uncheckedWeakness
GPT-5 agent misses included exhausted step budgets, repeated searches and early termination.