Artificial Analysis Legal Index
A combined score from seven evaluations, weighted to reflect the skills used in legal work.
Reported results
| Claude Fable 5.1 (max with fallback)Anthropic | 60 | |
|---|---|---|
| GPT-6 Astra (max)OpenAI | 59 | |
| Claude Fable 5 (with fallback)Anthropic | 58 | |
| Claude Opus 5 (max)Anthropic | 57 | |
| Grok 4.6 (high)xAI | 52 | |
| Gemini 3.8 Flash (high)Google | 52 | |
| Muse Spark 1.3 (max)Harvey | 51 | |
| GPT-5.6 Sol (max)OpenAI | 50 | |
| Kimi K3 (max)Moonshot AI | 49 | |
| Qwen3.8 Max (0902)Alibaba | 46 |
Only the knowledge and non-hallucination components explicitly test law. The other evaluations measure general abilities and are weighted to reflect legal work. The live chart shows 60 for its leader while the page FAQ mentions an unlisted 61-point configuration, so this snapshot records the chart.
What the benchmark measures
A combined score from seven evaluations, weighted to reflect the skills used in legal work.
The evaluated unit is a model response or component output. Read the source for the exact prompt, tool and harness conditions.
How it is scored
Weighted composite: legal knowledge 35%, agentic knowledge work 25%, reasoning 15%, long context 10%, non-hallucination 10%, tool use 5%.
Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.
Sources
- Primary benchmark source
- Artificial Analysis live Legal Index chart · results 2026-09 · checked 2026-09-18
Useful for a broad comparison, but several components test general abilities rather than legal work. The chart and FAQ name different leading configurations; the table here follows the chart.