RedlineBench
RedlineBench contains 140 tasks across three simulated SaaS and professional-services negotiations. Agents produce native Word redlines and comments over four negotiation turns.
Reported results
| GPT-5.5 | 50.5 | |
|---|---|---|
| Claude Fable 5 | 47.3 | |
| Gemini 3.5 Flash | 45.1 | |
| Claude Opus 4.8 | 44.4 |
Fable 5 was tested once; the other systems were tested three times. Each of the twelve combinations of scenario and turn has equal weight in the overall score.
RedlineBench, opening round30.3
| GPT-5.5 | 30.3 | |
|---|---|---|
| Claude Fable 5 | 22.6 | |
| Gemini 3.5 Flash | 21.9 | |
| Claude Opus 4.8 | 17.9 |
The first redline against a counterparty draft, before any back-and-forth. Every system scores far lower here than on its overall figure.
What the benchmark measures
140 tasks involving edits to documents across three SaaS and services negotiations, each with four rounds.
The evaluated unit is an agent or completed task. Read the source for the exact prompt, tool and harness conditions.
How it is scored
Validity gate, weighted attorney rubrics, three-judge majority vote.
Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.
Sources
- Primary benchmark source
- Crosby Intelligence, Jun 2026 · results 2026-06 · checked 2026-09-18
- Crosby Intelligence, Jun 2026 · results 2026-06 · checked 2026-09-18
Tests DOCX redlining across negotiation rounds, including whether the edits are valid.