GPT-5.5, medium reasoning · Results from 2026-08
ContractScrub
Contracts prepared by lawyers with realistic errors introduced for a final review, including broken references and inconsistent defined terms.
Reported results
What the benchmark measures
Contracts prepared by lawyers with realistic errors introduced for a final review, including broken references and inconsistent defined terms.
The evaluated unit is a model response or component output. Read the source for the exact prompt, tool and harness conditions.
How it is scored
Macro recall by error class.
Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.
Sources
Tests the 'last pair of eyes' review. Error types drawn from real malpractice claims.