MultiChallenge
Multi-turn conversations testing whether a model keeps instructions, remembers what was inferred, stays consistent with itself and edits earlier versions reliably.
Reported results
| Claude 3.5 Sonnet (Jun 2024) | 58.57 | |
|---|---|---|
| o1-preview | 34.29 | |
| Gemini 1.5 Pro (Aug 2024) | 31.43 | |
| Mistral Large | 21.43 | |
| GPT-4o (Aug 2024) | 14.29 | |
| Llama 3.1 405B Instruct | 12.86 |
These are older models, tested on examples selected because they were difficult for them. The results help illustrate the failure mode; they do not rank current systems.
MultiChallenge, inference memory41.53
| o1-preview | 41.53 | |
|---|---|---|
| Claude 3.5 Sonnet (Jun 2024) | 37.29 | |
| Llama 3.1 405B Instruct | 16.95 | |
| Gemini 1.5 Pro (Aug 2024) | 15.25 | |
| Mistral Large | 9.32 | |
| GPT-4o (Aug 2024) | 5.08 |
Whether the model remembers something it worked out earlier in the conversation, rather than something it was told. Same historical cohort.
MultiChallenge, versioned editing39.02
| o1-preview | 39.02 | |
|---|---|---|
| Claude 3.5 Sonnet (Jun 2024) | 24.39 | |
| Gemini 1.5 Pro (Aug 2024) | 19.51 | |
| GPT-4o (Aug 2024) | 17.07 | |
| Mistral Large | 7.32 | |
| Llama 3.1 405B Instruct | 4.88 |
Editing an earlier version of a text correctly after later turns have changed it. The lowest category for most models in the cohort.
MultiChallenge, self-coherence45.45
| Claude 3.5 Sonnet (Jun 2024) | 45.45 | |
|---|---|---|
| o1-preview | 34.09 | |
| Llama 3.1 405B Instruct | 25 | |
| Mistral Large | 20.45 | |
| GPT-4o (Aug 2024) | 13.64 | |
| Gemini 1.5 Pro (Aug 2024) | 13.64 |
Whether the model stays consistent with what it said earlier. Same historical cohort.
What the benchmark measures
Multi-turn conversations testing whether a model keeps instructions, remembers what was inferred, stays consistent with itself and edits earlier versions reliably.
The evaluated unit is a model response or component output. Read the source for the exact prompt, tool and harness conditions.
How it is scored
Human-evaluated success rate per challenge category.
Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.
Sources
- Primary benchmark source
- MultiChallenge paper, Table 2, Jan 2025 · results 2025-01 · checked 2026-09-18
- MultiChallenge paper, Table 2, Jan 2025 · results 2025-01 · checked 2026-09-18
- MultiChallenge paper, Table 2, Jan 2025 · results 2025-01 · checked 2026-09-18
- MultiChallenge paper, Table 2, Jan 2025 · results 2025-01 · checked 2026-09-18
Tests general conversation skills relevant to negotiation and drafting. Its separate categories help examine where continuity breaks down.