Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

Legal Research Bench

Research · Agent or completed task · Open · US federal and state · 2026

Legal Research Bench evaluates tool-using agents on questions involving statutes, regulations and case law across US federal and state jurisdictions. The published harness includes web search, page parsing, stored-document retrieval and CourtListener search.

Reported results

Questions passing every required rubric item · higher is betterResults 2026-09 · checked 2026-09-18
Legal Research Bench, strict completion
Muse Spark 1.3 MaxHarvey
55.29
Claude Opus 5Anthropic
55.29
Claude Fable 5.1Anthropic
55.29
Claude Fable 5Anthropic
49.52
GLM-5.3Zhipu AI
49.04
Grok 4.6xAI
48.08
GPT-5.6 SolOpenAI
48.08
Qwen 3.8 MaxAlibaba
47.6

The top three tie at 55.29%. Claude Opus 5 reaches 90.58% under weighted partial credit but only 55.29% when every required item must pass. Conflicting-authority questions reduce scores by 6–17 points per model.

Top eight of 61 reported systems; US federal and state research across eight practice areas. Source: VALS AI, updated 15 Sep 2026.

What the benchmark measures

Legal research questions across eight practice areas that require agents to find and combine sources. Answers are assessed for substance and supporting authority.

The evaluated unit is an agent or completed task. Read the source for the exact prompt, tool and harness conditions.

How it is scored

Two measures: questions passing every required item, and weighted scores that award partial credit.

Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.

Sources

Under the strict measure, every required item must pass. The public leaderboard is updated separately from the original open-source release.