Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

ContractEval

Contract · Model or response · Open · Commercial · 2025

Clause-level risk questions derived from CUAD, tested on open and proprietary models.

Reported results

What the benchmark measures

Clause-level risk questions derived from CUAD, tested on open and proprietary models.

The evaluated unit is a model response or component output. Read the source for the exact prompt, tool and harness conditions.

How it is scored

Correctness and output effectiveness scores.

Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.

Sources

Extends CUAD from extraction into explanation.