Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

Artificial Analysis Legal Index

Reasoning · Model or response · Mixed · Cross-jurisdictional composite · 2026

A combined score from seven evaluations, weighted to reflect the skills used in legal work.

Reported results

Occupationally weighted composite score · higher is betterResults 2026-09 · checked 2026-09-18
Artificial Analysis Legal Index
Claude Fable 5.1 (max with fallback)Anthropic
60
GPT-6 Astra (max)OpenAI
59
Claude Fable 5 (with fallback)Anthropic
58
Claude Opus 5 (max)Anthropic
57
Grok 4.6 (high)xAI
52
Gemini 3.8 Flash (high)Google
52
Muse Spark 1.3 (max)Harvey
51
GPT-5.6 Sol (max)OpenAI
50
Kimi K3 (max)Moonshot AI
49
Qwen3.8 Max (0902)Alibaba
46

Only the knowledge and non-hallucination components explicitly test law. The other evaluations measure general abilities and are weighted to reflect legal work. The live chart shows 60 for its leader while the page FAQ mentions an unlisted 61-point configuration, so this snapshot records the chart.

Top ten of 24 displayed model configurations; seven component evaluations. Source: Artificial Analysis live Legal Index chart.

What the benchmark measures

A combined score from seven evaluations, weighted to reflect the skills used in legal work.

The evaluated unit is a model response or component output. Read the source for the exact prompt, tool and harness conditions.

How it is scored

Weighted composite: legal knowledge 35%, agentic knowledge work 25%, reasoning 15%, long context 10%, non-hallucination 10%, tool use 5%.

Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.

Sources

Useful for a broad comparison, but several components test general abilities rather than legal work. The chart and FAQ name different leading configurations; the table here follows the chart.