Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

LongBench v2

Documents · Model or response · Open · Domain-general · 2024

503 questions using source material ranging from 8,000 to 2 million words.

Reported results

Overall multiple-choice accuracy, with chain of thought · higher is betterResults 2025-07 · checked 2026-09-18
LongBench v2, overall
Gemini 2.5 Pro
63.3
Gemini 2.5 Flash
62.1
Qwen3-235B-A22B Thinking (Jul 2025)
60.6
DeepSeek-R1
58.3
Qwen3-235B-A22B Instruct (Jul 2025)
58.3

Human experts score 53.7 on the same questions. The overall figure mixes document types and does not isolate legal precedence, scanned pages or multi-document questions.

Top five of the official leaderboard's chain-of-thought column, rows dated January to July 2025. Source: LongBench v2 official leaderboard, THUDM.

What the benchmark measures

503 questions using source material ranging from 8,000 to 2 million words.

The evaluated unit is a model response or component output. Read the source for the exact prompt, tool and harness conditions.

How it is scored

Multiple-choice accuracy by context length.

Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.

Sources

Tests whether models actually use long contexts.