Harvey Legal Agent Benchmark
Harvey’s Legal Agent Benchmark contains more than 1,200 client-matter tasks across 24 practice areas. Each task includes instructions, a file-based matter environment and a required reviewable work product.
Reported results
| Muse Spark 1.2Harvey | 25.42 | |
|---|---|---|
| Muse Spark 1.3 MaxHarvey | 23.75 | |
| Muse Spark 1.3Harvey | 22.08 | |
| Muse Spark 1.1Harvey | 20 | |
| Grok 4.6xAI | 15.83 | |
| Claude Fable 5Anthropic | 11.25 |
Individual criterion pass rates are high, around 92–95%, but strict full-task resolution reaches only 20–25% for the leading systems. Passing most checks can still leave the task unfinished.
What the benchmark measures
More than 1,200 tasks spanning multiple steps across 24 practice areas, with over 75,000 assessment criteria.
The evaluated unit is an agent or completed task. Read the source for the exact prompt, tool and harness conditions.
How it is scored
Criterion pass rate and all-criteria task completion.
Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.
Sources
Tests sequences of legal work involving tools, research and drafting.