Tasks
- Contract review
- Legal research
- Advisory
- Negotiation
- Due diligence
- Drafting
- Obligations and matter management
ScoredMapped, no score—Nothing recordedGapWeaknessSelect any cell to open it
Contract review
| Step | ContractEvalmodel or response | ContractNLImodel or response | ContractScrubmodel or response | CUADmodel or response | RedlineBenchagent, completed task |
|---|---|---|---|---|---|
| Clause span extraction | — | — | — | 44Precision at 80% recall | — |
| Entailment and contradiction | — | Mapped | — | — | — |
| Absence detectionGap | — | Mapped | — | — | — |
| Risk identification | Mapped | — | — | — | — |
| Defined terms and cross-references | — | — | Mapped | — | — |
| Overall redlining | — | — | — | — | 50.5Turn-weighted rubric score |
Legal research
| Step | LegalBench-RAGmodel or response | LePhantomCitemodel or response | RAGBenchmodel or response | Legal Research Benchagent, completed task |
|---|---|---|---|---|
| Retrieval and evidence spans | Mapped | — | — | — |
| Citation | — | — | — | Mapped |
| Citation verificationWeakness | — | 68.8Agentic citation-verification F1 | — | — |
| Answer faithfulness | — | — | Mapped | — |
| Legal research, every requirement metWeakness | — | — | — | 55.29Questions passing every required rubric item |
| Current authority checksGap | — | — | — | Mapped |
Advisory
| Step | HELM Enterprise (Legal)model or response | LawBenchmodel or response | LegalBenchmodel or response | PLawBenchmodel or response | PRBench Legalmodel or response | Realm: Legalagent, completed task |
|---|---|---|---|---|---|---|
| Issue spotting | Mapped | — | Mapped | — | — | Mapped |
| Rule recall | — | Mapped | Mapped | — | — | — |
| Rule application | — | — | Mapped | Mapped | — | Mapped |
| Interpretation | — | Mapped | Mapped | — | — | — |
| Professional reasoning | — | — | — | — | Mapped | — |
| Reasoning through a changing matterWeakness | — | — | — | — | — | 62.1Mean weighted rubric score |
| Uncertainty and escalationGap | — | — | — | — | Mapped | — |
Negotiation
| Step | MultiChallengemodel or response | PLawBenchmodel or response | NegotiationArenaagent, completed task | RedlineBenchagent, completed task |
|---|---|---|---|---|
| Overall redlining | — | — | — | 50.5Turn-weighted rubric score |
| Opening-round redlining | — | — | — | 30.3Turn 1 rubric score |
| Legal correctness | — | Mapped | — | Mapped |
| Commercial alignmentGap | — | — | — | Mapped |
| Mandate and concession consistency | — | — | Mapped | Mapped |
| Self-coherenceGap | 45.45Human-evaluated success | — | — | — |
Due diligence
| Step | ContractEvalmodel or response | ContractScrubmodel or response | CUADmodel or response | LongBench v2model or response | MAUDmodel or response | MP-DocVQAmodel or response |
|---|---|---|---|---|---|---|
| Deal point extraction | — | — | — | — | Mapped | — |
| Clause span extraction | — | — | 44Precision at 80% recall | — | — | — |
| Multi-document questions | — | — | — | Mapped | — | — |
| Visual and scanned documents | — | — | — | — | — | Mapped |
| Document precedenceGap | — | Mapped | — | Mapped | — | — |
| Risk identification | Mapped | — | — | — | — | — |
Drafting
| Step | ContractScrubmodel or response | HELM Enterprise (Legal)model or response | LexSummmodel or response | MultiChallengemodel or response | PLawBenchmodel or response | Harvey LABagent, completed task | RedlineBenchagent, completed task |
|---|---|---|---|---|---|---|---|
| Drafting | — | — | — | — | Mapped | Mapped | — |
| Overall redlining | — | — | — | — | — | — | 50.5Turn-weighted rubric score |
| Versioned editing | — | — | — | 39.02Human-evaluated success | — | — | — |
| Defined terms and cross-references | Mapped | — | — | — | — | — | — |
| Legal correctness | — | — | — | — | Mapped | — | Mapped |
| Summarisation | — | Mapped | Mapped | — | — | — | — |
Obligations and matter management
| Step | MultiChallengemodel or response | Harvey LABagent, completed task | Tau Benchagent, completed task |
|---|---|---|---|
| Instruction retention | 58.57Human-evaluated success | — | — |
| Inference memory | 41.53Human-evaluated success | — | — |
| Policy compliance | — | — | Mapped |
| Legal agent tasks, all criteriaWeakness | — | 25.42Tasks passing every rubric criterion | — |
| Application reproducibilityGap | — | — | Mapped |
Records
CUAD · tableCUAD, precision at 80% recall44
Precision at 80% recall, span-level · higher is better · results 2021-03 · checked 2026-09-18
Ten fine-tuned models from the paper; original test split.
Historical baselines from 2021, not a current ranking. This measure asks how many extracted passages are correct when the model finds 80% of the relevant passages.
| System | Score |
|---|---|
| DeBERTa-xlarge | 44 |
| RoBERTa-large | 38.1 |
| RoBERTa-base, contracts pretraining | 34.1 |
| RoBERTa-base | 31.1 |
| ALBERT-xxlarge | 31 |
| ALBERT-large | 20.9 |
| ALBERT-xlarge | 20.5 |
| ALBERT-base | 11.1 |
| BERT-base | 8.2 |
| BERT-large | 7.6 |
RedlineBench · tableRedlineBench, overall50.5
Turn-weighted rubric score, validity gate applied · higher is better · results 2026-06 · checked 2026-09-18
Four systems, three SaaS negotiation scenarios, four turns each.
Fable 5 was tested once; the other systems were tested three times. Each of the twelve combinations of scenario and turn has equal weight in the overall score.
| System | Score |
|---|---|
| GPT-5.5 | 50.5 |
| Claude Fable 5 | 47.3 |
| Gemini 3.5 Flash | 45.1 |
| Claude Opus 4.8 | 44.4 |
LePhantomCite · tableLePhantomCite, citation verification68.8
Agentic citation-verification F1 · higher is better · results 2026-06 · checked 2026-09-18
Six agentic systems evaluated on 1,300 legal brief excerpts with injected citation problems.
F1 balances precision and recall when identifying citation problems. The paper also studies generated citations over time. Those results are separate from this verification table.
| System | Note | Score |
|---|---|---|
| Claude Code, Opus 4.8 | P 76.1 · R 62.8 | 68.8 |
| GPT-5 | P 40.8 · R 84.4 | 55 |
| Qwen3.6-27B | P 28.9 · R 65.0 | 40 |
| GPT-OSS-120B | P 21.1 · R 55.1 | 30.5 |
| Gemini 2.5 Flash | P 16.9 · R 66.9 | 27 |
| Qwen3-8B | P 12.0 · R 41.1 | 18.6 |
Legal Research Bench · tableLegal Research Bench, strict completion55.29
Questions passing every required rubric item · higher is better · results 2026-09 · checked 2026-09-18
Top eight of 61 reported systems; US federal and state research across eight practice areas.
The top three tie at 55.29%. Claude Opus 5 reaches 90.58% under weighted partial credit but only 55.29% when every required item must pass. Conflicting-authority questions reduce scores by 6–17 points per model.
| System | Note | Score |
|---|---|---|
| Muse Spark 1.3 Max | Harvey | 55.29 |
| Claude Opus 5 | Anthropic | 55.29 |
| Claude Fable 5.1 | Anthropic | 55.29 |
| Claude Fable 5 | Anthropic | 49.52 |
| GLM-5.3 | Zhipu AI | 49.04 |
| Grok 4.6 | xAI | 48.08 |
| GPT-5.6 Sol | OpenAI | 48.08 |
| Qwen 3.8 Max | Alibaba | 47.6 |
Realm: Legal · tableRealm: Legal, long-horizon reasoning62.1
Mean weighted rubric score · higher is better · results 2026-09 · checked 2026-09-18
The ten highest-scoring configurations from 16 reported; litigation, transactional and compliance tasks where the record changes.
The leaderboard includes newer models than the original three-model analysis. That earlier report found problems with selecting rules, applying facts, recognising missing information and revising later work. Those findings should not be attributed to the newer models without separate testing.
| System | Note | Score |
|---|---|---|
| Claude Opus 5 (max) | Anthropic | 62.1 |
| Claude Fable 5.1 (max) | Anthropic | 60.8 |
| Claude Fable 5 | Anthropic | 55.7 |
| Kimi K3 (max) | Moonshot AI | 54.1 |
| Grok 4.6 (high) | xAI | 53.8 |
| GPT-5.6 Sol (max) | OpenAI | 51.1 |
| Muse Spark 1.3 (xhigh) | Harvey | 50.3 |
| Gemini 3.8 Flash (high) | 47.9 | |
| Grok 4.5 (high) | xAI | 43.2 |
| Muse Spark 1.1 (xhigh) | Harvey | 42.3 |
RedlineBench · tableRedlineBench, opening round30.3
Turn 1 rubric score · higher is better · results 2026-06 · checked 2026-09-18
Four systems, first turn only.
The first redline against a counterparty draft, before any back-and-forth. Every system scores far lower here than on its overall figure.
| System | Score |
|---|---|
| GPT-5.5 | 30.3 |
| Claude Fable 5 | 22.6 |
| Gemini 3.5 Flash | 21.9 |
| Claude Opus 4.8 | 17.9 |
MultiChallenge · tableMultiChallenge, self-coherence45.45
Human-evaluated success · higher is better · results 2025-01 · checked 2026-09-18
Six models from 2024, general conversation.
Whether the model stays consistent with what it said earlier. Same historical cohort.
| System | Score |
|---|---|
| Claude 3.5 Sonnet (Jun 2024) | 45.45 |
| o1-preview | 34.09 |
| Llama 3.1 405B Instruct | 25 |
| Mistral Large | 20.45 |
| GPT-4o (Aug 2024) | 13.64 |
| Gemini 1.5 Pro (Aug 2024) | 13.64 |
MultiChallenge · tableMultiChallenge, versioned editing39.02
Human-evaluated success · higher is better · results 2025-01 · checked 2026-09-18
Six models from 2024, general conversation.
Editing an earlier version of a text correctly after later turns have changed it. The lowest category for most models in the cohort.
| System | Score |
|---|---|
| o1-preview | 39.02 |
| Claude 3.5 Sonnet (Jun 2024) | 24.39 |
| Gemini 1.5 Pro (Aug 2024) | 19.51 |
| GPT-4o (Aug 2024) | 17.07 |
| Mistral Large | 7.32 |
| Llama 3.1 405B Instruct | 4.88 |
MultiChallenge · tableMultiChallenge, instruction retention58.57
Human-evaluated success · higher is better · results 2025-01 · checked 2026-09-18
Six models from 2024, general conversation.
These are older models, tested on examples selected because they were difficult for them. The results help illustrate the failure mode; they do not rank current systems.
| System | Score |
|---|---|
| Claude 3.5 Sonnet (Jun 2024) | 58.57 |
| o1-preview | 34.29 |
| Gemini 1.5 Pro (Aug 2024) | 31.43 |
| Mistral Large | 21.43 |
| GPT-4o (Aug 2024) | 14.29 |
| Llama 3.1 405B Instruct | 12.86 |
MultiChallenge · tableMultiChallenge, inference memory41.53
Human-evaluated success · higher is better · results 2025-01 · checked 2026-09-18
Six models from 2024, general conversation.
Whether the model remembers something it worked out earlier in the conversation, rather than something it was told. Same historical cohort.
| System | Score |
|---|---|
| o1-preview | 41.53 |
| Claude 3.5 Sonnet (Jun 2024) | 37.29 |
| Llama 3.1 405B Instruct | 16.95 |
| Gemini 1.5 Pro (Aug 2024) | 15.25 |
| Mistral Large | 9.32 |
| GPT-4o (Aug 2024) | 5.08 |
Harvey Legal Agent Benchmark · tableHarvey LAB (legal agent tasks)25.42
Tasks passing every rubric criterion · higher is better · results 2026-09 · checked 2026-09-18
Six current systems; a task passes only if every criterion passes.
Individual criterion pass rates are high, around 92–95%, but strict full-task resolution reaches only 20–25% for the leading systems. Passing most checks can still leave the task unfinished.
| System | Note | Score |
|---|---|---|
| Muse Spark 1.2 | Harvey | 25.42 |
| Muse Spark 1.3 Max | Harvey | 23.75 |
| Muse Spark 1.3 | Harvey | 22.08 |
| Muse Spark 1.1 | Harvey | 20 |
| Grok 4.6 | xAI | 15.83 |
| Claude Fable 5 | Anthropic | 11.25 |
Benchmark · model or response · CommercialContractEval64.4
Clause-level risk questions derived from CUAD, tested on open and proprietary models.
Scored by
Correctness and output effectiveness scores.
Headline result
64.4 Correctness F1 on the CUAD test set
Supports
Performance on its clause-level risk questions and output-effectiveness criteria.
Does not establish
A complete review of an agreement or fitness for a particular organisation’s risk position.
Extends CUAD from extraction into explanation.
Benchmark · model or response · Commercial NDAsContractNLI89.2
607 contracts tested against policy-like hypotheses, with evidence spans for each judgement.
Scored by
Three-way classification accuracy and evidence identification F1.
Headline result
89.2 NLI accuracy, macro-averaged over hypotheses
Supports
Classification and supporting-span identification for its contract hypotheses.
Does not establish
Open-ended review, negotiation or advice against a client’s playbook.
Tests whether a contract entails, contradicts, or is silent on a policy statement.
Benchmark · model or response · CommercialContractScrub75
Contracts prepared by lawyers with realistic errors introduced for a final review, including broken references and inconsistent defined terms.
Scored by
Macro recall by error class.
Headline result
75 Macro-average recall over nine error categories
Supports
Detection of the error classes deliberately introduced into its lawyer-prepared contracts.
Does not establish
Detection of every drafting defect or substantive legal issue in an unseen contract.
Tests the 'last pair of eyes' review. Error types drawn from real malpractice claims.
Benchmark · model or response · US commercialCUAD44
510 contracts with 41 clause types and 13,000+ expert annotations for clause-level extraction.
Scored by
Span-level AUPR and precision at fixed recall.
Headline result
44 Precision at 80% recall, span-level
Supports
Clause identification and extraction performance across its 510 contracts and 41 clause categories.
Does not establish
Legal advice, contract risk judgment or complete contract review.
An established test of contract extraction. The results recorded here are historical baselines.
Benchmark · agent, completed task · US commercialRedlineBench50.5
140 tasks involving edits to documents across three SaaS and services negotiations, each with four rounds.
Scored by
Validity gate, weighted attorney rubrics, three-judge majority vote.
Headline result
50.5 Turn-weighted rubric score, validity gate applied
Supports
Performance on its multi-turn contract negotiations, document mechanics and attorney-authored rubrics.
Does not establish
Other agreement types, governing laws, playbooks, parties, bargaining positions or commercial contexts.
Tests DOCX redlining across negotiation rounds, including whether the edits are valid.
Benchmark · model or response · English legalLegalBench-RAG
6,858 expert-annotated query-answer pairs with character-level evidence spans.
Scored by
Character-level precision and recall.
Headline result
No comparable table recorded.
Supports
Retrieval and evidence-span performance on its annotated legal query-answer pairs.
Does not establish
Complete legal research, authority validation or a supported client-facing conclusion.
Tests retrieval and grounding at character level.
Benchmark · model or response · US litigationLePhantomCite68.8
1,300 legal brief excerpts containing deliberately introduced citation problems.
Scored by
Precision and recall by hallucination category.
Headline result
68.8 Agentic citation-verification F1
Supports
Detection of the citation defects represented in its legal-brief excerpts.
Does not establish
General factual accuracy, complete authority checking or substantive correctness of the brief.
Tests whether systems can identify fabricated or otherwise problematic citations.
Benchmark · model or response · Domain-generalRAGBench
Labelled examples for assessing retrieval-augmented generation (RAG): systems that retrieve source material before generating an answer.
Scored by
TRACe metrics for retrieval and generation quality.
Headline result
No comparable table recorded.
Supports
Retrieval and generation quality on its labelled RAG examples and TRACe measures.
Does not establish
Legal-domain accuracy or sufficiency for a particular legal research task.
Covers general domains; its evaluation approach can inform legal retrieval systems.
Benchmark · agent, completed task · US federal and stateLegal Research Bench55.29
Legal research questions across eight practice areas that require agents to find and combine sources. Answers are assessed for substance and supporting authority.
Scored by
Two measures: questions passing every required item, and weighted scores that award partial credit.
Headline result
55.29 Questions passing every required rubric item
Supports
Performance on its legal research questions using the supplied tools, sources, agent configuration and scoring method.
Does not establish
Other jurisdictions, an organisation’s research sources, completeness against a live matter record or reliability of downstream client advice.
Under the strict measure, every required item must pass. The public leaderboard is updated separately from the original open-source release.
Benchmark · model or response · Primarily USHELM Enterprise (Legal)
IBM extension of Stanford HELM with legal-specific scenarios.
Scored by
HELM 7-metric framework.
Headline result
No comparable table recorded.
Supports
Performance on the legal scenarios and metrics included in the HELM extension.
Does not establish
Complete legal-service quality or performance outside those scenarios.
Assesses legal scenarios using several measures of performance.
Benchmark · model or response · Chinese legalLawBench56.3
20 tasks across three cognitive levels following Bloom's taxonomy.
Scored by
Task-specific accuracy across cognitive levels.
Headline result
56.3 Average score over 20 tasks, zero-shot
Supports
Performance across its 20 Chinese-law tasks and three cognitive levels.
Does not establish
Cross-jurisdictional legal capability or end-to-end legal work.
Published at EMNLP 2024; covers a range of Chinese legal tasks.
Benchmark · model or response · Primarily USLegalBench88.6
162 tasks contributed by legal professionals, covering six types of legal reasoning.
Scored by
Task-specific exact match and classification accuracy.
Headline result
88.6 Mean accuracy across 162 tasks
Supports
Performance on its defined mixture of legal-language and rule tasks under the recorded prompting and scoring conditions.
Does not establish
An end-to-end client matter, current legal research, completeness of a fact record, advice against client objectives, sustained matter state or safe production action.
Broad coverage of legal reasoning types. Also used in Stanford HELM.
Benchmark · model or response · Chinese legalPLawBench69.7
Assesses language models on legal practice tasks using detailed scoring criteria.
Scored by
Rubric-guided scores across practice dimensions.
Headline result
69.7 Overall rubric score, weighted across three tasks
Supports
Performance on its Chinese legal-practice tasks and rubric dimensions.
Does not establish
Performance in other jurisdictions or in a particular live practice environment.
Uses assessment rubrics intended to reflect how lawyers evaluate work.
Benchmark · model or response · Multiple jurisdictions; published geographic totals cover legal and financeProfessional Reasoning Benchmark — Legal
500 legal questions developed with professionals, including a harder subset of 250. Tasks can involve several turns and use 10–30 weighted assessment criteria.
Scored by
A model judge applies weighted criteria to produce a score bounded between 0 and 1. The source also reports category analysis and confidence intervals.
Headline result
No comparable table recorded.
Supports
Performance on its professional legal questions and weighted assessment criteria.
Does not establish
Complete matter performance, current research or operation in a legal-service environment.
The displayed leaderboard does not specify whether it covers the full legal set or the harder subset. The discussion of the harder subset refers to older models. No overall score is recorded here while that distinction remains unclear.
Benchmark · agent, completed task · US federal and stateRealm: Legal62.1
Litigation, transactional and compliance tasks completed over several steps, using legal materials and tools as the facts or authority change.
Scored by
Mean score across 35–60 weighted criteria per task, organised around issue, rule, application and conclusion (IRAC). The source also reports the best of three attempts.
Headline result
62.1 Mean weighted rubric score
Supports
Performance on changing, long-horizon legal scenarios within the supplied sandbox, tools and rubric.
Does not establish
Performance with a firm’s own matter files, legal sources, templates, reviewers, permissions or production systems.
The recorded leaderboard covers 16 model configurations. Its detailed failure analysis concerns the original three models, so the IRAC breakdown should not be applied to newer entries.
Benchmark · model or response · Domain-generalMultiChallenge
Multi-turn conversations testing whether a model keeps instructions, remembers what was inferred, stays consistent with itself and edits earlier versions reliably.
Scored by
Human-evaluated success rate per challenge category.
Headline result
No comparable table recorded.
Supports
Performance on its instruction-retention, inference-memory, consistency and version-editing challenges.
Does not establish
Legal accuracy or reliable matter-state management in production.
Tests general conversation skills relevant to negotiation and drafting. Its separate categories help examine where continuity breaks down.
Benchmark · agent, completed task · Domain-generalNegotiationArena
Negotiation environments covering bargaining and resource exchange between language-model agents.
Scored by
Scenario-specific negotiation outcomes.
Headline result
No comparable table recorded.
Supports
Outcomes within its bargaining and resource-exchange environments.
Does not establish
Legal negotiation quality, client alignment or performance in document-based negotiations.
Tests strategic behaviour in general negotiation settings. It does not establish whether an agent follows a legal mandate.
Benchmark · model or response · Domain-generalLongBench v263.3
503 questions using source material ranging from 8,000 to 2 million words.
Scored by
Multiple-choice accuracy by context length.
Headline result
63.3 Overall multiple-choice accuracy, with chain of thought
Supports
Question-answering performance over its long-context source materials.
Does not establish
Legal reasoning, source authority or dependable use of a live matter record.
Tests whether models actually use long contexts.
Benchmark · model or response · US public M&AMAUD57.8
152 merger agreements with 92 questions and 47,000+ labels across material deal points.
Scored by
Minority-class AUPR per question.
Headline result
57.8 Mean minority-class AUPR over all questions, single-task
Supports
Identification of labelled merger-agreement deal points represented in the dataset.
Does not establish
Complete M&A review, transaction strategy or advice on an unseen deal.
Tests extraction of deal points from merger agreements, making it relevant to M&A diligence.
Benchmark · model or response · Domain-generalMP-DocVQA88.2
Questions over multi-page scanned industry documents.
Scored by
Answer exact match and page retrieval accuracy.
Headline result
88.2 ANLS on the hidden test set
Supports
Answer and page-retrieval performance over its scanned multi-page documents.
Does not establish
Legal interpretation, document completeness or reliable OCR on every document type.
Relevant to reading scanned legal documents, though the benchmark itself is not law-specific.
Benchmark · model or response · US, UK, EU, IndiaLexSumm
Eight legal summarisation datasets across multiple jurisdictions.
Scored by
Generation quality metrics requiring legal-domain supplements.
Headline result
No comparable table recorded.
Supports
Summarisation performance on its eight legal datasets under the reported metrics.
Does not establish
Accuracy for every audience, purpose, jurisdiction or downstream legal use.
Useful for comparing legal summarisation across jurisdictions.
Benchmark · agent, completed task · Primarily USHarvey Legal Agent Benchmark25.42
More than 1,200 tasks spanning multiple steps across 24 practice areas, with over 75,000 assessment criteria.
Scored by
Criterion pass rate and all-criteria task completion.
Headline result
25.42 Tasks passing every rubric criterion
Supports
Performance on its long-horizon, file-based legal assignments and expert criteria under the stated harness conditions.
Does not establish
Complete coverage of legal practice or reliable operation in a particular firm’s information, supervision and production environment.
Tests sequences of legal work involving tools, research and drafting.
Benchmark · agent, completed task · Domain-generalTau Bench69.2
Conversations where an agent uses domain-specific tools and must follow the relevant policies.
Scored by
Task success rate and policy compliance score.
Headline result
69.2 pass^1 task success, retail domain
Supports
Task success and policy compliance in its conversational tool-use environments.
Does not establish
Legal correctness, legal authority to act or compliance with a legal team’s specific policy.
Tests whether agents follow rules while executing tasks.
Gap · partial support: ContractNLI, Realm: LegalMissing provisions and factsGap
Spotting what a document does not contain.
A review can correctly assess every clause it finds and still miss the protection that was left out. Finding what should be there needs its own test.
ContractNLI uses a NotMentioned label when a contract is silent on a fixed statement. Realm: Legal’s original diagnostic analysis also found failures when important facts were missing. Neither establishes that a system can check your full document set against a list of required provisions.
Gap · no benchmark in the collectionIs the authority still good law?Gap
Whether cited authority is still good law.
A case can be real, relevant and accurately quoted, yet no longer be good law. Checking that a citation exists does not settle whether you can rely on it.
LePhantomCite checks citations and their support for a claim. Realm: Legal tests revisions when new authority or a changed timeframe is supplied. Those tests do not establish that a system can independently discover that a case has been overruled or a provision amended.
Gap · partial support: Professional Reasoning Benchmark — LegalKnowing when to answer or escalateGap
Whether the system recognises uncertainty and knows when to answer, ask for missing information or hand the question to a person.
A confident wrong answer can be more dangerous than no answer. Adding “possibly” to the sentence does little to help the person relying on it.
Professional Reasoning Benchmark — Legal (PRBench) assesses how models handle uncertainty within its expert rubrics. That provides some evidence, but it does not separately test the workflow decision to answer, request information or refer the matter to a person.
Gap · no benchmark in the collectionCommercial judgementGap
Whether a clause is acceptable for this deal, given its value, duration, counterparty and the client’s priorities.
The same clause can carry very different risks in a short pilot and a ten-year outsourcing deal. A playbook cannot settle every trade-off.
The benchmarks here do not directly test when commercial context justifies departing from a standard position. RedlineBench applies a playbook, which covers part of the work but leaves that judgement open.
Gap · no benchmark in the collectionHolding a position under challengeGap
Whether the system changes a supported analysis just because the user pushes back.
“Are you sure?” should prompt a check of the reasoning. It should not be enough, on its own, to reverse the conclusion.
No benchmark reviewed here directly isolates pushback without new evidence. RedlineBench includes several turns, but those turns involve negotiation and counter-proposals.
Gap · no benchmark in the collectionWhich document governs?Gap
Whether the system resolves the relationship between the master agreement, amendments, side letters and schedules.
The answer may be in the third amendment, not the master agreement. Reading one file can give you a clear answer to the wrong version of the deal.
The contract tests reviewed here do not establish reliable precedence across a complete agreement set. LongBench v2 covers general reasoning across documents, but it does not isolate legal amendment chains or order-of-precedence clauses.
Gap · partial support: Tau BenchRun-to-run consistencyGap
Whether repeating the same task with the same documents and prompt produces materially consistent findings.
If two reviews of the same contract disagree, you need to know why. Otherwise it is hard to explain which findings someone should rely on.
Tau Bench measures success across repeated attempts at retail and airline tasks, using pass^k. That is useful background for legal systems, which need their own repeated-run tests. ContractScrub uses “consistency” differently: it checks terms within a document.
Weakness · LePhantomCite, reported agent runs.Valid citations flagged as falseWeakness
Some tested agents flagged valid citations that were absent from CourtListener. A missing lookup became a false alarm.
Weakness · LePhantomCite, GPT-5 with BOED.Verification stopped with items uncheckedWeakness
GPT-5 agent misses included exhausted step budgets, repeated searches and early termination.
Weakness · Harvey LAB and Legal Research Bench, VALS AI, September 2026.Partial success does not aggregate to completionWeakness
Harvey LAB: around 94.5% of individual criteria pass, 25.4% of tasks complete. Legal Research Bench: 90.58% under weighted credit, 55.29% when every required item must pass.
Weakness · Realm: Legal, micro1.Failures recognising missing information and revising later workWeakness
Realm: Legal’s original three-model analysis found problems selecting rules, applying facts, recognising missing information and revising later work. The analysis has not been repeated on newer entries.