Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

Gaps

CoverageNoneNo direct benchmark test.PartialSome of it is directly tested.ProxyOnly adjacent benchmark coverage.

StatusGapnobody has measured itWeaknessmeasured, and poorMeasuredthere is meaningful direct benchmark evidencebothwhere the failure lives

Contract review

7 gaps · 1 none · 2 partial · 4 proxy

Select a row for design and testing options.

Contract review: gaps down the rows, with the benchmark coverage recorded for each.
GapCoverageOpen
Run-to-run consistencyPartial
Missing provisions and factsPartial
Privilege and confidentialityNone
Carries context between stagesProxy
Follows policy while actingProxy
Completes repeatedly, not onceProxy
Knows when to stop or hand overProxy

Select a row for design and testing options.

Agentic problems

11 properties · 7 gaps · 3 weaknesses · 1 measured

Agentic problems down the rows, with where the failure lives and the status of each.
Agentic problemFailure lives inStatusOpen
Holds state across turnsbothWeakness
Carries context between stagesharnessGap
Revises when facts or authority changebothWeakness
Selects tools and builds argumentsmodelMeasured
Follows policy while actingbothGap
Completes repeatedly, not onceharnessGap
Completes the whole task, every criterionbothWeakness
Recovers from a failed stepharnessGap
Knows when to stop or hand overbothGap
Changes the system of record correctlyharnessGap
Acts within authority and permissionsharnessGap
Commercial judgementNone

A legally correct answer is not necessarily the right commercial answer for this transaction.

Solution designTest
  • Record the commercial context alongside the legal issue: objectives, priorities, risk appetite, leverage and acceptable fallbacks.
  • Use a playbook or decision framework that separates the legal position from the commercial choice.
  • Present meaningful trade-offs and options rather than forcing a single “correct” recommendation.
  • Send decisions with no clear rule or objectively correct answer to the appropriate human decision-maker.
  • Hold the legal issue constant and vary the commercial context.
  • Check whether recommendations change when priorities, leverage or risk appetite change.
  • Test whether the system surfaces the trade-off rather than resolving it silently.
  • Check that judgement-dependent decisions are handed over at the right point.

Editorial. Suggested, not tested.

Related tasks

AdvisoryNegotiation

Knowing when to answer or escalatePartial

A system can produce a plausible answer even when it does not have enough information to support one.

Solution designTest
  • Give the system clear routes for answer, ask, qualify, escalate or stop.
  • Require missing or conflicting information to be identified before a firm conclusion is given.
  • Keep uncertainty visible rather than converting it into confident prose.
  • Define categories of advice or action that always require human review.
  • Use cases with sufficient, insufficient and conflicting information.
  • Check whether missing information triggers a question rather than an invented assumption.
  • Measure both failures to escalate and unnecessary escalation.
  • Test whether the system chooses the right route, not just whether the final wording looks plausible.

Editorial. Suggested, not tested.

Related tasks

Advisory

Benchmark context

Professional Reasoning Benchmark — Legal · adjacent

View Professional Reasoning Benchmark — Legal →

Run-to-run consistencyPartial

The same workflow can produce materially different conclusions or actions across repeated runs.

Solution designTest
  • Keep one reliable record of the matter’s facts, instructions and previous decisions rather than relying only on the conversation.
  • Use fixed rules for things that should not vary, such as permissions, routing, calculations and completion checks.
  • Capture important conclusions and proposed actions in structured fields rather than only in free text.
  • Check important outputs before they trigger an action or update.
  • Keep human approval for decisions where variation is acceptable but the consequence of acting on it is material.
  • Run the same matter several times under the same conditions.
  • Compare the legal conclusion and proposed action, not just differences in wording.
  • Repeat the test after changing the model, prompt, sources or workflow.
  • Identify where materially different outcomes first arise.

Editorial. Suggested, not tested.

Related tasks

Contract reviewLegal researchObligations & matter managementCross-task evaluation

Benchmark context

Tau Bench · adjacent, domain-general agent benchmark

View Tau Bench →

Missing provisions and factsPartial

A system can correctly assess everything it finds while failing to notice something that should have been there.

Solution designTest
  • Define what information or provisions are expected before reviewing what is actually present.
  • Record clear states such as present, absent, not found, not searched and unclear.
  • Keep extraction and completeness checking as separate steps.
  • Record what documents, sources or sections were actually searched.
  • Remove expected clauses, facts or documents and check whether the omission is identified.
  • Test whether “not present” is distinguished from “not found”.
  • Include cases where every extracted item is correct but the overall answer is incomplete.
  • Vary the available document set and check whether the system recognises the limits of its search.

Editorial. Suggested, not tested.

Related tasks

Contract reviewLegal researchDue diligence

Benchmark context

ContractNLI · adjacentRealm: Legal · adjacent

View ContractNLI →
View Realm: Legal →

Is the authority still good law?None

Finding a real citation does not establish that it remains authoritative for the question being answered.

Solution designTest
  • Record the jurisdiction, date, legal status and proposition supported for each important authority.
  • Treat citation verification and current-law checking as separate tasks.
  • Define which sources and levels of authority are acceptable for the work.
  • Flag authorities that are outdated, superseded, materially limited or not yet checked.
  • Use authorities whose later treatment differs.
  • Include valid citations that are nevertheless outdated or superseded.
  • Change the relevant jurisdiction or date and check whether the answer changes appropriately.
  • Verify that both the citation and its current status were checked.

Editorial. Suggested, not tested.

Related tasks

Legal research

Holding a position under challengeNone

A system may change a correct conclusion because a user challenges it rather than because the underlying facts or law changed.

Solution designTest
  • Keep the conclusion together with the facts, sources and reasoning that support it.
  • Treat disagreement as a reason to review the conclusion, not automatically reverse it.
  • Require a recorded reason whenever a material conclusion changes.
  • Distinguish new information from pressure, tone or repetition.
  • Challenge a correct conclusion repeatedly without adding new information.
  • Change tone and confidence while keeping the underlying facts fixed.
  • Measure unjustified reversals.
  • Then introduce genuinely new information and check that the conclusion changes when it should.

Editorial. Suggested, not tested.

Related tasks

Negotiation

Which document governs?None

Correct analysis of the wrong document is still the wrong answer.

Solution designTest
  • Record the relationship between the main agreement, amendments, schedules, orders and other related documents.
  • Identify dates, supersession and precedence before substantive analysis begins.
  • Use the current governing version as the basis for downstream work.
  • Keep an explicit unresolved state where the document hierarchy is genuinely unclear.
  • Create document sets with different amendment and precedence structures.
  • Ask the same question against different document histories.
  • Remove a relevant amendment and check whether uncertainty is surfaced.
  • Test conflicting and genuinely ambiguous document relationships.

Editorial. Suggested, not tested.

Related tasks

Due diligence

Bias in reviews and scoringPartial

Variation can enter through both the system producing the answer and the system judging it.

Solution designTest
  • Use a clear scoring framework with separate criteria for different aspects of quality.
  • Keep irrelevant characteristics out of the scoring process where possible.
  • Use an independent second review for high-impact or disputed assessments.
  • Keep scoring rules and evaluator versions controlled so changes are visible.
  • Hold the substance constant while varying irrelevant characteristics.
  • Compare results across relevant groups, drafting styles or scenarios.
  • Compare automated scoring with blinded human review.
  • Run the same outputs through different evaluators and check whether the result materially changes.

Editorial. Suggested, not tested.

Related tasks

Cross-task evaluation

Benchmark context

FairLex · adjacent

View FairLex →

Privilege and confidentialityNone

A substantively correct answer can still be a failure if information is retrieved, combined or disclosed improperly.

Solution designTest
  • Classify information by matter, user, sensitivity and permitted use.
  • Give the system access only to information the user and matter are entitled to use.
  • Prevent restricted material from being added to prompts or outputs simply because it is technically retrievable.
  • Apply disclosure checks before information is sent, shared or used in an action.
  • Seed environments with material carrying different access restrictions.
  • Test cross-user and cross-matter retrieval.
  • Check whether restricted information enters prompts or outputs.
  • Test correct-answer / wrong-recipient scenarios.

Editorial. Suggested, not tested.

Related tasks

Contract reviewDue diligence

Who is the rationale for?None

The explanation for an edit should change with its audience: an internal note, client advice and a message to the counterparty are not the same document.

Solution designTest
  • Make the intended audience explicit for every explanation.
  • Keep the substantive edit separate from the explanation so the same change can support different rationales.
  • Define what may and may not be disclosed to each audience, including strategy and fallback positions.
  • Treat tone and style as separate from the underlying legal position.
  • Hold the edit constant and vary the audience.
  • Check that the substantive edit remains the same where it should.
  • Test whether internal reasoning or fallback positions leak into counterparty-facing text.
  • Compare explanations across audiences and check that only the framing, not the legal position, changes.

Editorial. Suggested, not tested.

Related tasks

Drafting

Carries context between stagesProxy

Facts, decisions and constraints established earlier in a matter may disappear or become distorted later in the workflow.

Solution designTest
  • Keep a shared matter record containing the current facts, decisions, assumptions and unresolved issues.
  • Have each stage read from and update that same record rather than relying on conversation history alone.
  • Record material changes so later work can see what changed and why.
  • Give each stage the matter information it needs without passing irrelevant history forward.
  • Use multi-stage work where information established early matters later.
  • Check whether key constraints survive transitions between stages.
  • Change an earlier fact and check whether dependent work updates.
  • Resume a workflow at a later stage and compare the result with a continuous run.

Editorial. Suggested, not tested.

HARNESS

Related tasks

Contract reviewLegal researchAdvisoryNegotiationDue diligenceObligations & matter management

Benchmark context

Realm: Legal · partialHarvey Legal Agent Benchmark · partial

View Realm: Legal →
View Harvey Legal Agent Benchmark →

Follows policy while actingProxy

An agent can complete the requested action while still breaching the organisation’s rules for how that action should be performed.

Solution designTest
  • Keep organisational rules separate from ordinary user instructions.
  • Turn important policies into explicit checks the system must pass before acting.
  • Block or escalate actions that conflict with policy rather than relying on the model to remember the rule.
  • Record why a consequential action was allowed, blocked or escalated.
  • Create tasks where the requested action conflicts with policy.
  • Vary the wording while keeping the policy rule constant.
  • Attempt otherwise valid actions that fail a policy condition.
  • Check whether the reason an action was allowed or blocked can be reviewed afterwards.

Editorial. Suggested, not tested.

BOTH

Related tasks

Contract reviewAdvisoryNegotiationObligations & matter management

Benchmark context

Tau Bench · direct, domain-general agent benchmark

View Tau Bench →

Completes repeatedly, not onceProxy

One successful run does not establish that a workflow will complete reliably in production.

Solution designTest
  • Define what complete means for the whole workflow, including every mandatory step.
  • Keep the workflow’s current state explicit so incomplete and blocked work cannot be mistaken for finished work.
  • Make consequential actions safe to retry without duplicating them.
  • Check the final state before the workflow is marked complete.
  • Repeat the same workflow multiple times.
  • Measure whole-workflow completion rather than average step success.
  • Run repeated trials after changes to the model, prompt, tools or workflow.
  • Check whether every required step succeeds on every run.

Editorial. Suggested, not tested.

HARNESS

Related tasks

Contract reviewLegal researchDue diligenceObligations & matter managementCross-task evaluation

Benchmark context

RedlineBench · partialTau Bench · direct, domain-general agent benchmark

View RedlineBench →
View Tau Bench →

Recovers from a failed stepProxy

Long workflows will encounter unavailable tools, timeouts, malformed outputs and partial failures.

Solution designTest
  • Define which failures can be retried, which need an alternative route and which must stop for human help.
  • Save progress at important stages so work can resume safely.
  • Limit retries and provide a fallback rather than allowing repeated failure loops.
  • Provide a way to undo or correct partial actions where necessary.
  • Inject tool failures, timeouts and malformed responses.
  • Check whether retries resolve the issue rather than duplicate it.
  • Interrupt execution and resume from a saved point.
  • Test partial completion and whether the workflow returns to a valid state.

Editorial. Suggested, not tested.

HARNESS

Related tasks

Legal researchDue diligenceObligations & matter management

Knows when to stop or hand overProxy

Continuing to act can be as dangerous as failing to act.

Solution designTest
  • Define clear stop and escalation conditions as part of the workflow.
  • Distinguish complete, blocked, outside authority and needs review from an ordinary answer.
  • Specify who receives the matter when it is handed over.
  • Pass the human reviewer the relevant facts, work completed and unresolved issues rather than making them reconstruct the history.
  • Create cases where continuing would be inappropriate.
  • Check whether the system stops at the correct point.
  • Test incomplete, blocked and out-of-authority cases.
  • Check whether a human can continue without rebuilding the matter from scratch.

Editorial. Suggested, not tested.

BOTH

Related tasks

Contract reviewLegal researchAdvisoryNegotiation

Benchmark context

Professional Reasoning Benchmark — Legal · adjacent

View Professional Reasoning Benchmark — Legal →

Changes the system of record correctlyProxy

A correct recommendation does not mean the correct record, field or status was actually changed.

Solution designTest
  • Limit the system to a defined set of permitted records, fields and actions.
  • Check the target record and proposed change before writing it.
  • Require approval for higher-impact changes.
  • Confirm after the action that the intended record was actually updated.
  • Compare the expected and actual system state after execution.
  • Test wrong-record and wrong-field scenarios.
  • Test invalid, incomplete and conflicting updates.
  • Test duplicate actions, partial writes and failed persistence.

Editorial. Suggested, not tested.

HARNESS

Related tasks

Obligations & matter management

Acts within authority and permissionsProxy

An agent may be technically capable of performing an action that the user, matter or workflow does not authorise.

Solution designTest
  • Give the system only the access it needs for the current task.
  • Separate permission to recommend, draft and execute.
  • Check authority immediately before consequential actions.
  • Require human approval where the requested action exceeds ordinary delegated authority.
  • Attempt actions just outside the user’s permitted scope.
  • Test escalation from permitted read access to prohibited write access.
  • Check whether approval is requested at the correct point.
  • Give permission at one level and test whether the system exceeds it.

Editorial. Suggested, not tested.

HARNESS

Related tasks

AdvisoryNegotiationObligations & matter management

Holds state across turnsWeakness

Important facts, instructions, positions or document changes can be lost as a conversation gets longer.

Solution designTest
  • Keep a persistent matter record outside the conversation.
  • Version important instructions, positions, findings and document changes.
  • Make the current matter record, not chat history, the source of truth for later steps.
  • Record which document or version is current before further work continues.
  • Ask the system to restate its current position after a long interaction.
  • Compare that position with the matter record.
  • Change an earlier instruction and check whether later work uses the updated version.
  • Test whether superseded information reappears.

Editorial. Suggested, not tested.

BOTH

Benchmark context

MultiChallenge · adjacent, domain-general model benchmark

View MultiChallenge →

Revises when facts or authority changeWeakness

When a material fact or authority changes, the system may fail to revisit conclusions that depended on it.

Solution designTest
  • Record which conclusions depend on which facts and authorities.
  • When an input changes, flag the conclusions that may be affected.
  • Revisit affected work rather than rerunning everything or ignoring the change.
  • Show what changed and why when a conclusion is revised.
  • Change one material fact late in the workflow.
  • Replace or supersede a relied-on authority.
  • Check whether every affected conclusion is revisited.
  • Confirm that unrelated conclusions remain stable.

Editorial. Suggested, not tested.

BOTH

Benchmark context

Realm: Legal · direct

View Realm: Legal →

Selects tools and builds argumentsMeasured

The system must choose the right tool or source for the task and use the result to support a coherent legal argument.

Solution designTest
  • Define which tools and sources are appropriate for each type of task.
  • Separate research and retrieval from the legal reasoning that uses the material.
  • Require important propositions to be tied back to the source or tool output that supports them.
  • Provide a clear fallback where a required tool is unavailable or returns an uncertain result.
  • Give the system tasks requiring different tools and sources.
  • Check whether it chooses the appropriate route.
  • Test whether the argument actually uses the material retrieved.
  • Introduce unavailable or conflicting tool results and check how the system responds.

Editorial. Suggested, not tested.

MODEL

Benchmark context

Legal Research Bench · partialLegalAgentBench · directBFCL v4 · direct, domain-general agent benchmark

View Legal Research Bench →
View LegalAgentBench →
View BFCL v4 →

Completes the whole task, every criterionWeakness

Strong performance on most individual requirements can still leave the overall legal task unfinished.

Solution designTest
  • Define every mandatory requirement for the task before the system begins.
  • Keep unmet or unresolved requirements visible throughout the workflow.
  • Do not allow a polished final answer to hide missing work.
  • Use a final completion check before the task can be marked finished.
  • Score both individual requirements and whole-task completion.
  • Require every mandatory item to pass.
  • Remove one required item from an otherwise strong result and check whether the task remains incomplete.
  • Track which requirements most often prevent completion.

Editorial. Suggested, not tested.

BOTH

Benchmark context

Harvey Legal Agent Benchmark · directLegal Research Bench · directRedlineBench · direct

View Harvey Legal Agent Benchmark →
View Legal Research Bench →
View RedlineBench →