LongBench v2
503 questions using source material ranging from 8,000 to 2 million words.
Reported results
Overall multiple-choice accuracy, with chain of thought · higher is better
| Gemini 2.5 Pro | 63.3 | |
|---|---|---|
| Gemini 2.5 Flash | 62.1 | |
| Qwen3-235B-A22B Thinking (Jul 2025) | 60.6 | |
| DeepSeek-R1 | 58.3 | |
| Qwen3-235B-A22B Instruct (Jul 2025) | 58.3 |
Human experts score 53.7 on the same questions. The overall figure mixes document types and does not isolate legal precedence, scanned pages or multi-document questions.
What the benchmark measures
503 questions using source material ranging from 8,000 to 2 million words.
The evaluated unit is a model response or component output. Read the source for the exact prompt, tool and harness conditions.
How it is scored
Multiple-choice accuracy by context length.
Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.
Sources
- Primary benchmark source
- LongBench v2 official leaderboard, THUDM · results 2025-07 · checked 2026-09-18
Tests whether models actually use long contexts.