Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

Weaknesses

Measured, with results that expose a practical weakness.

Select a row for design and testing options.

Outcomes down the rows, with what was recorded and which benchmark recorded it.
OutcomeResultBenchmarkOpen
Reasoning through a changing matterAdvisoryFailures recognising missing information and revising later workRealm: Legal
Clause span extractionContract review44Precision at 80% recall48.2Span-level AUPRCUAD
Versioned editingDrafting39.02Human-evaluated successMultiChallenge
Citation verificationLegal researchValid citations flagged as falseVerification stopped with items uncheckedLePhantomCite
Whole-task legal researchLegal researchPartial success does not aggregate to completionHarvey Legal Agent Benchmark
Legal Research Bench
Opening-round redliningNegotiation30.3Turn 1 rubric scoreRedlineBench
Self-coherenceNegotiation45.45Human-evaluated successMultiChallenge
Inference memoryObligations and matter management41.53Human-evaluated successMultiChallenge
Whole legal-agent taskObligations and matter management94.52%25.42%individual criteria passed → whole tasks resolvedHarvey Legal Agent Benchmark
Reliability across repeated runsAgentic workflows4622.5single-run success → successful in all four runsairline domain69.246.2single-run success → successful in all four runsretail domainTau BenchDomain-general agent benchmark
Reasoning through a changing matter
Solution designTest
  • Keep a structured matter record outside the conversation so facts, assumptions and decisions remain clear as the matter changes.
  • Record which conclusions depend on which facts or authorities.
  • When something material changes, identify and revisit only the conclusions affected.
  • Keep a clear history of what changed and why.
  • Change one material fact after the initial work is produced.
  • Replace or supersede a relied-on authority.
  • Change the relevant jurisdiction or date.
  • Check whether affected conclusions are revised and unaffected conclusions stay stable.

Editorial. Suggested, not tested.

BOTH

Related gap

Revises when facts or authority change

Benchmark context

Realm: Legal · direct

View Realm: Legal →

Clause span extraction
Solution designTest
  • Keep extraction and completeness checking as separate steps.
  • Capture clauses in a consistent structure rather than as isolated free text.
  • Keep each extracted item linked to its location in the source document.
  • Check for expected provisions that may be missing.
  • Distinguish “not present” from “not found”.
  • Measure precision and recall separately.
  • Include unusual drafting and clauses spread across several sections.
  • Include documents where expected provisions are deliberately absent.
  • Test downstream review when extraction is incomplete.
  • Verify extracted information against the source text.

Editorial. Suggested, not tested.

MODEL + HARNESS

Related gap

Missing provisions and facts

Benchmark context

CUAD · direct

View CUAD →

Versioned editing
Solution designTest
  • Keep a clear current version of the document rather than relying on conversation history.
  • Record edits against that version.
  • Preserve accepted, rejected and superseded changes.
  • Make later edits against the current document, not an earlier copy.
  • Use stable references to provisions so changes do not depend only on paragraph position.
  • Apply several sequential edits.
  • Change an earlier instruction after later edits have been made.
  • Test conflicting and overlapping amendments.
  • Compare the final document with the expected final version.
  • Check for lost, duplicated or reintroduced text.

Editorial. Suggested, not tested.

BOTH

Related gap

Carries context between stages

Benchmark context

MultiChallenge · direct, domain-general model benchmark

View MultiChallenge →

Citation verification
Solution designTest
  • Treat four questions separately: does the citation exist, is it accurate, does it support the proposition, and is it still current?
  • Keep verification results independently reviewable from the generated answer.
  • Allow unable to verify rather than forcing every citation into valid or invalid.
  • Keep every citation in a clear checked / failed / unresolved state.
  • Do not allow unchecked citations to disappear from the workflow.
  • Mix valid, fabricated, miscited and superseded authorities.
  • Measure false positives as well as missed bad citations.
  • Include partial verification runs and tool failures.
  • Check that every citation ends in an explicit state.
  • Compare final quality with and without the verification step.

Editorial. Suggested, not tested.

BOTH

Related gap

Is the authority still good law?

Benchmark context

LePhantomCite · direct

View LePhantomCite →

Whole-task legal research
Solution designTest
  • Define the mandatory requirements for the research task before work begins.
  • Separate essential requirements from optional quality improvements.
  • Keep unresolved questions, missing authorities and unchecked propositions visible.
  • Require a completion check before the work can be treated as finished.
  • Keep important propositions linked to the sources that support them.
  • Score individual requirements and whole-task completion separately.
  • Require every mandatory requirement to pass.
  • Include questions requiring several authorities or propositions.
  • Remove one required item from an otherwise strong answer.
  • Record which missing requirement prevented completion.

Editorial. Suggested, not tested.

HARNESS

Related gap

Completes the whole task, every criterion

Benchmark context

Harvey Legal Agent Benchmark · directLegal Research Bench · direct

View Harvey Legal Agent Benchmark →
View Legal Research Bench →

Opening-round redlining
Solution designTest
  • Separate three steps: identify the issue, decide the negotiation position, then draft the change.
  • Ground proposed edits in the relevant playbook or stated position.
  • Keep the reason for each material change.
  • Track unresolved points between negotiation rounds.
  • Check whether an edit creates inconsistency elsewhere in the agreement.
  • Test first-pass redlining separately from later-round improvement.
  • Include interacting clauses and defined terms.
  • Compare proposed changes against an explicit playbook.
  • Measure issue coverage as well as drafting quality.
  • Check whether fixing one provision causes a problem elsewhere.

Editorial. Suggested, not tested.

BOTH

Related gap

Commercial judgement

Benchmark context

RedlineBench · direct

View RedlineBench →

Self-coherence
Solution designTest
  • Keep important conclusions together with the reasons and material that support them.
  • Treat user disagreement as a prompt to review, not a reason to change position automatically.
  • Require a clear reason for material changes in conclusion.
  • Re-check dependent conclusions after a justified revision.
  • Keep important constraints in the matter record rather than relying on conversational memory.
  • Challenge a correct conclusion repeatedly without changing the facts.
  • Change tone and phrasing while keeping the substance fixed.
  • Measure unjustified reversals.
  • Then introduce genuinely new information.
  • Check that the conclusion changes when, and only when, the new information warrants it.

Editorial. Suggested, not tested.

MODEL + HARNESS

Related gap

Holding a position under challenge

Benchmark context

MultiChallenge · direct, domain-general model benchmark

View MultiChallenge →

Inference memory
Solution designTest
  • Keep important derived conclusions in the matter record rather than only in the conversation.
  • Record which facts or decisions each conclusion depends on.
  • Distinguish source facts from conclusions drawn from them.
  • Give later steps only the matter information relevant to the current task.
  • When an underlying fact changes, identify and revisit conclusions that depended on it.
  • Use multi-stage work where an early conclusion matters much later.
  • Resume the work after several intervening steps.
  • Change an underlying fact and check whether the old conclusion is discarded.
  • Compare conversation-only memory with a structured matter record.
  • Test whether unrelated earlier information affects later reasoning.

Editorial. Suggested, not tested.

BOTH

Related gap

Carries context between stages

Benchmark context

MultiChallenge · direct, domain-general model benchmark

View MultiChallenge →

Whole legal-agent task
Solution designTest
  • Define every mandatory requirement for the complete task.
  • Keep incomplete, partial and complete as distinct states.
  • Preserve unresolved requirements between steps.
  • Require a final completion check before the task can be marked finished.
  • Treat “mostly correct” and “finished” as different outcomes.
  • Report individual requirement success and whole-task completion separately.
  • Require every mandatory requirement to pass.
  • Record which requirement caused an otherwise strong run to remain incomplete.
  • Repeat representative end-to-end tasks.
  • Re-test after changes to the model, prompt, tools or workflow.

Editorial. Suggested, not tested.

HARNESS

Related gap

Completes the whole task, every criterion

Benchmark context

Harvey Legal Agent Benchmark · direct

View Harvey Legal Agent Benchmark →

Reliability across repeated runs
Solution designTest
  • Keep one reliable record of the workflow’s facts, instructions and current state.
  • Use fixed rules for decisions that should not vary, such as permissions, routing and completion checks.
  • Capture important conclusions and actions in a structured form rather than only in free text.
  • Check important outputs before they trigger consequential actions.
  • Make repeated actions safe so a retry does not create duplicates or inconsistent records.
  • Repeat the same representative workflow several times.
  • Report single-run success and repeated-success rates separately.
  • Measure whether all required steps succeed on every run.
  • Identify which step causes degradation across repeated attempts.
  • Re-run the test after changes to the model, prompt, tools or workflow.

Editorial. Suggested, not tested.

HARNESS

Related gap

Completes repeatedly, not once

Benchmark context

Tau Bench · adjacent, domain-general agent benchmark · airline domain, retail domain

View Tau Bench →