LegalBench-RAG
6,858 expert-annotated query-answer pairs with character-level evidence spans.
Reported results
No comparable score table is recorded in this collection.
What the benchmark measures
6,858 expert-annotated query-answer pairs with character-level evidence spans.
The evaluated unit is a model response or component output. Read the source for the exact prompt, tool and harness conditions.
How it is scored
Character-level precision and recall.
Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.