Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

RedlineBench

Contract · Agent or completed task · Open · US commercial · 2025

RedlineBench contains 140 tasks across three simulated SaaS and professional-services negotiations. Agents produce native Word redlines and comments over four negotiation turns.

Reported results

Turn-weighted rubric score, validity gate applied · higher is betterResults 2026-06 · checked 2026-09-18
RedlineBench, overall
GPT-5.5
50.5
Claude Fable 5
47.3
Gemini 3.5 Flash
45.1
Claude Opus 4.8
44.4

Fable 5 was tested once; the other systems were tested three times. Each of the twelve combinations of scenario and turn has equal weight in the overall score.

Four systems, three SaaS negotiation scenarios, four turns each. Source: Crosby Intelligence, Jun 2026.

RedlineBench, opening round30.3
Turn 1 rubric score · higher is betterResults 2026-06 · checked 2026-09-18
RedlineBench, opening round
GPT-5.5
30.3
Claude Fable 5
22.6
Gemini 3.5 Flash
21.9
Claude Opus 4.8
17.9

The first redline against a counterparty draft, before any back-and-forth. Every system scores far lower here than on its overall figure.

Four systems, first turn only. Source: Crosby Intelligence, Jun 2026.

What the benchmark measures

140 tasks involving edits to documents across three SaaS and services negotiations, each with four rounds.

The evaluated unit is an agent or completed task. Read the source for the exact prompt, tool and harness conditions.

How it is scored

Validity gate, weighted attorney rubrics, three-judge majority vote.

Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.

Sources

Tests DOCX redlining across negotiation rounds, including whether the edits are valid.