Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

Professional Reasoning Benchmark — Legal

Reasoning · Model or response · Mixed · Multiple jurisdictions; published geographic totals cover legal and finance · 2026

500 legal questions developed with professionals, including a harder subset of 250. Tasks can involve several turns and use 10–30 weighted assessment criteria.

Reported results

No comparable score table is recorded in this collection.

What the benchmark measures

500 legal questions developed with professionals, including a harder subset of 250. Tasks can involve several turns and use 10–30 weighted assessment criteria.

The evaluated unit is a model response or component output. Read the source for the exact prompt, tool and harness conditions.

How it is scored

A model judge applies weighted criteria to produce a score bounded between 0 and 1. The source also reports category analysis and confidence intervals.

Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.

Sources

The displayed leaderboard does not specify whether it covers the full legal set or the harder subset. The discussion of the harder subset refers to older models. No overall score is recorded here while that distinction remains unclear.