Span NLI BERT (DeBERTa-v2-xlarge backbone) · Results from 2021-10
ContractNLI
607 contracts tested against policy-like hypotheses, with evidence spans for each judgement.
Reported results
What the benchmark measures
607 contracts tested against policy-like hypotheses, with evidence spans for each judgement.
The evaluated unit is a model response or component output. Read the source for the exact prompt, tool and harness conditions.
How it is scored
Three-way classification accuracy and evidence identification F1.
Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.
Sources
Tests whether a contract entails, contradicts, or is silent on a policy statement.