Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

CUAD

Contract · Model or response · Open · US commercial · 2021

510 contracts with 41 clause types and 13,000+ expert annotations for clause-level extraction.

Reported results

Precision at 80% recall, span-level · higher is betterResults 2021-03 · checked 2026-09-18
CUAD, precision at 80% recall
DeBERTa-xlarge
44
RoBERTa-large
38.1
RoBERTa-base, contracts pretraining
34.1
RoBERTa-base
31.1
ALBERT-xxlarge
31
ALBERT-large
20.9
ALBERT-xlarge
20.5
ALBERT-base
11.1
BERT-base
8.2
BERT-large
7.6

Historical baselines from 2021, not a current ranking. This measure asks how many extracted passages are correct when the model finds 80% of the relevant passages.

Ten fine-tuned models from the paper; original test split. Source: Hendrycks et al., CUAD paper, Table 2.

CUAD, area under the precision-recall curve48.2
Span-level AUPR · higher is betterResults 2021-03 · checked 2026-09-18
CUAD, area under the precision-recall curve
RoBERTa-large
48.2
DeBERTa-xlarge
47.8
RoBERTa-base, contracts pretraining
45.2
RoBERTa-base
42.6
ALBERT-xxlarge
38.4
ALBERT-xlarge
37.8
ALBERT-base
35.3
ALBERT-large
34.9
BERT-base
32.4
BERT-large
32.3

Summarises the whole precision-recall curve rather than one operating point. RoBERTa-large scores slightly higher than DeBERTa on this measure.

Same ten models as the precision table. Source: Hendrycks et al., CUAD paper, Table 2.

What the benchmark measures

510 contracts with 41 clause types and 13,000+ expert annotations for clause-level extraction.

The evaluated unit is a model response or component output. Read the source for the exact prompt, tool and harness conditions.

How it is scored

Span-level AUPR and precision at fixed recall.

Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.

Sources

An established test of contract extraction. The results recorded here are historical baselines.