Weaknesses
Measured, with results that expose a practical weakness.
Select a row for design and testing options.
| Outcome | Result | Benchmark | Open |
|---|---|---|---|
| Reasoning through a changing matterAdvisory | Failures recognising missing information and revising later work | Realm: Legal | |
| Clause span extractionContract review | 44Precision at 80% recall48.2Span-level AUPR | CUAD | |
| Versioned editingDrafting | 39.02Human-evaluated success | MultiChallenge | |
| Citation verificationLegal research | Valid citations flagged as falseVerification stopped with items unchecked | LePhantomCite | |
| Whole-task legal researchLegal research | Partial success does not aggregate to completion | Harvey Legal Agent Benchmark Legal Research Bench | |
| Opening-round redliningNegotiation | 30.3Turn 1 rubric score | RedlineBench | |
| Self-coherenceNegotiation | 45.45Human-evaluated success | MultiChallenge | |
| Inference memoryObligations and matter management | 41.53Human-evaluated success | MultiChallenge | |
| Whole legal-agent taskObligations and matter management | 94.52%→25.42%individual criteria passed → whole tasks resolved | Harvey Legal Agent Benchmark | |
| Reliability across repeated runsAgentic workflows | 46→22.5single-run success → successful in all four runsairline domain69.2→46.2single-run success → successful in all four runsretail domain | Tau BenchDomain-general agent benchmark |
Reasoning through a changing matter
| Solution design | Test |
|---|---|
|
|
Editorial. Suggested, not tested.
BOTH
Related gap
Revises when facts or authority change
Benchmark context
Realm: Legal · direct
Clause span extraction
| Solution design | Test |
|---|---|
|
|
Editorial. Suggested, not tested.
MODEL + HARNESS
Related gap
Benchmark context
CUAD · direct
Versioned editing
| Solution design | Test |
|---|---|
|
|
Editorial. Suggested, not tested.
BOTH
Related gap
Carries context between stages
Benchmark context
MultiChallenge · direct, domain-general model benchmark
Citation verification
| Solution design | Test |
|---|---|
|
|
Editorial. Suggested, not tested.
BOTH
Related gap
Is the authority still good law?
Benchmark context
LePhantomCite · direct
Whole-task legal research
| Solution design | Test |
|---|---|
|
|
Editorial. Suggested, not tested.
HARNESS
Related gap
Completes the whole task, every criterion
Benchmark context
Harvey Legal Agent Benchmark · directLegal Research Bench · direct
View Harvey Legal Agent Benchmark →
View Legal Research Bench →
Opening-round redlining
| Solution design | Test |
|---|---|
|
|
Editorial. Suggested, not tested.
BOTH
Related gap
Benchmark context
RedlineBench · direct
Self-coherence
| Solution design | Test |
|---|---|
|
|
Editorial. Suggested, not tested.
MODEL + HARNESS
Related gap
Holding a position under challenge
Benchmark context
MultiChallenge · direct, domain-general model benchmark
Inference memory
| Solution design | Test |
|---|---|
|
|
Editorial. Suggested, not tested.
BOTH
Related gap
Carries context between stages
Benchmark context
MultiChallenge · direct, domain-general model benchmark
Whole legal-agent task
| Solution design | Test |
|---|---|
|
|
Editorial. Suggested, not tested.
HARNESS
Related gap
Completes the whole task, every criterion
Benchmark context
Harvey Legal Agent Benchmark · direct
Reliability across repeated runs
| Solution design | Test |
|---|---|
|
|
Editorial. Suggested, not tested.
HARNESS
Related gap
Completes repeatedly, not once
Benchmark context
Tau Bench · adjacent, domain-general agent benchmark · airline domain, retail domain