LegalBench
LegalBench is an open-science collection of 162 legal reasoning tasks contributed by legal practitioners, academics and computational-law researchers. Its six categories are issue spotting, rule recall, rule conclusion, rule application, interpretation and rhetorical understanding.
Reported results
| Claude Fable 5Anthropic | 88.6 | |
|---|---|---|
| GPT-5.5OpenAI | 87.4 | |
| Claude Opus 5Anthropic | 86.9 | |
| Gemini 2.5 ProGoogle | 86.2 | |
| Grok 4.3xAI | 85.1 | |
| GLM-5.3Zhipu AI | 84.8 | |
| DeepSeek V4 FlashDeepSeek | 82.3 |
There are 6.3 percentage points between first and last, and 3.8 across the top six. Small differences on this test offer limited guidance on how to build your system.
What the benchmark measures
162 tasks contributed by legal professionals, covering six types of legal reasoning.
The evaluated unit is a model response or component output. Read the source for the exact prompt, tool and harness conditions.
How it is scored
Task-specific exact match and classification accuracy.
Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.
Sources
Broad coverage of legal reasoning types. Also used in Stanford HELM.