Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

CLAUSE

Assurance · Model or response · Open · Commercial · 2025

More than 7,500 contracts with deliberate changes across ten categories of anomaly.

Reported results

No comparable score table is recorded in this collection.

What the benchmark measures

More than 7,500 contracts with deliberate changes across ten categories of anomaly.

The evaluated unit is a model response or component output. Read the source for the exact prompt, tool and harness conditions.

How it is scored

Detection accuracy and explanation quality.

Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.

Sources

Tests whether models find and explain deliberately introduced errors.