Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

MultiChallenge

Assurance · Model or response · Open · Domain-general · 2025

Multi-turn conversations testing whether a model keeps instructions, remembers what was inferred, stays consistent with itself and edits earlier versions reliably.

Reported results

Human-evaluated success · higher is betterResults 2025-01 · checked 2026-09-18
MultiChallenge, instruction retention
Claude 3.5 Sonnet (Jun 2024)
58.57
o1-preview
34.29
Gemini 1.5 Pro (Aug 2024)
31.43
Mistral Large
21.43
GPT-4o (Aug 2024)
14.29
Llama 3.1 405B Instruct
12.86

These are older models, tested on examples selected because they were difficult for them. The results help illustrate the failure mode; they do not rank current systems.

Six models from 2024, general conversation. Source: MultiChallenge paper, Table 2, Jan 2025.

MultiChallenge, inference memory41.53
Human-evaluated success · higher is betterResults 2025-01 · checked 2026-09-18
MultiChallenge, inference memory
o1-preview
41.53
Claude 3.5 Sonnet (Jun 2024)
37.29
Llama 3.1 405B Instruct
16.95
Gemini 1.5 Pro (Aug 2024)
15.25
Mistral Large
9.32
GPT-4o (Aug 2024)
5.08

Whether the model remembers something it worked out earlier in the conversation, rather than something it was told. Same historical cohort.

Six models from 2024, general conversation. Source: MultiChallenge paper, Table 2, Jan 2025.

MultiChallenge, versioned editing39.02
Human-evaluated success · higher is betterResults 2025-01 · checked 2026-09-18
MultiChallenge, versioned editing
o1-preview
39.02
Claude 3.5 Sonnet (Jun 2024)
24.39
Gemini 1.5 Pro (Aug 2024)
19.51
GPT-4o (Aug 2024)
17.07
Mistral Large
7.32
Llama 3.1 405B Instruct
4.88

Editing an earlier version of a text correctly after later turns have changed it. The lowest category for most models in the cohort.

Six models from 2024, general conversation. Source: MultiChallenge paper, Table 2, Jan 2025.

MultiChallenge, self-coherence45.45
Human-evaluated success · higher is betterResults 2025-01 · checked 2026-09-18
MultiChallenge, self-coherence
Claude 3.5 Sonnet (Jun 2024)
45.45
o1-preview
34.09
Llama 3.1 405B Instruct
25
Mistral Large
20.45
GPT-4o (Aug 2024)
13.64
Gemini 1.5 Pro (Aug 2024)
13.64

Whether the model stays consistent with what it said earlier. Same historical cohort.

Six models from 2024, general conversation. Source: MultiChallenge paper, Table 2, Jan 2025.

What the benchmark measures

Multi-turn conversations testing whether a model keeps instructions, remembers what was inferred, stays consistent with itself and edits earlier versions reliably.

The evaluated unit is a model response or component output. Read the source for the exact prompt, tool and harness conditions.

How it is scored

Human-evaluated success rate per challenge category.

Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.

Sources

Tests general conversation skills relevant to negotiation and drafting. Its separate categories help examine where continuity breaks down.