CUAD
510 contracts with 41 clause types and 13,000+ expert annotations for clause-level extraction.
Reported results
| DeBERTa-xlarge | 44 | |
|---|---|---|
| RoBERTa-large | 38.1 | |
| RoBERTa-base, contracts pretraining | 34.1 | |
| RoBERTa-base | 31.1 | |
| ALBERT-xxlarge | 31 | |
| ALBERT-large | 20.9 | |
| ALBERT-xlarge | 20.5 | |
| ALBERT-base | 11.1 | |
| BERT-base | 8.2 | |
| BERT-large | 7.6 |
Historical baselines from 2021, not a current ranking. This measure asks how many extracted passages are correct when the model finds 80% of the relevant passages.
CUAD, area under the precision-recall curve48.2
| RoBERTa-large | 48.2 | |
|---|---|---|
| DeBERTa-xlarge | 47.8 | |
| RoBERTa-base, contracts pretraining | 45.2 | |
| RoBERTa-base | 42.6 | |
| ALBERT-xxlarge | 38.4 | |
| ALBERT-xlarge | 37.8 | |
| ALBERT-base | 35.3 | |
| ALBERT-large | 34.9 | |
| BERT-base | 32.4 | |
| BERT-large | 32.3 |
Summarises the whole precision-recall curve rather than one operating point. RoBERTa-large scores slightly higher than DeBERTa on this measure.
What the benchmark measures
510 contracts with 41 clause types and 13,000+ expert annotations for clause-level extraction.
The evaluated unit is a model response or component output. Read the source for the exact prompt, tool and harness conditions.
How it is scored
Span-level AUPR and precision at fixed recall.
Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.
Sources
- Primary benchmark source
- Hendrycks et al., CUAD paper, Table 2 · results 2021-03 · checked 2026-09-18
- Hendrycks et al., CUAD paper, Table 2 · results 2021-03 · checked 2026-09-18
An established test of contract extraction. The results recorded here are historical baselines.