Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

ContractNLI

Contract · Model or response · Open · Commercial NDAs · 2021

607 contracts tested against policy-like hypotheses, with evidence spans for each judgement.

Reported results

What the benchmark measures

607 contracts tested against policy-like hypotheses, with evidence spans for each judgement.

The evaluated unit is a model response or component output. Read the source for the exact prompt, tool and harness conditions.

How it is scored

Three-way classification accuracy and evidence identification F1.

Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.

Sources

Tests whether a contract entails, contradicts, or is silent on a policy statement.