Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

Realm: Legal

Agent · Agent or completed task · Mixed · US federal and state · 2026

Realm: Legal uses multi-stage litigation, transactional and compliance tasks. A system produces a legal work product and revises it as facts, arguments, authority, jurisdiction or timeframe change.

Reported results

Mean weighted rubric score · higher is betterResults 2026-09 · checked 2026-09-18
Realm: Legal, long-horizon reasoning
Claude Opus 5 (max)Anthropic
62.1
Claude Fable 5.1 (max)Anthropic
60.8
Claude Fable 5Anthropic
55.7
Kimi K3 (max)Moonshot AI
54.1
Grok 4.6 (high)xAI
53.8
GPT-5.6 Sol (max)OpenAI
51.1
Muse Spark 1.3 (xhigh)Harvey
50.3
Gemini 3.8 Flash (high)Google
47.9
Grok 4.5 (high)xAI
43.2
Muse Spark 1.1 (xhigh)Harvey
42.3

The leaderboard includes newer models than the original three-model analysis. That earlier report found problems with selecting rules, applying facts, recognising missing information and revising later work. Those findings should not be attributed to the newer models without separate testing.

The ten highest-scoring configurations from 16 reported; litigation, transactional and compliance tasks where the record changes. Source: micro1 Realm: Legal, updated 2 Sep 2026.

What the benchmark measures

Litigation, transactional and compliance tasks completed over several steps, using legal materials and tools as the facts or authority change.

The evaluated unit is an agent or completed task. Read the source for the exact prompt, tool and harness conditions.

How it is scored

Mean score across 35–60 weighted criteria per task, organised around issue, rule, application and conclusion (IRAC). The source also reports the best of three attempts.

Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.

Sources

The recorded leaderboard covers 16 model configurations. Its detailed failure analysis concerns the original three models, so the IRAC breakdown should not be applied to newer entries.