Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

LegalBench-RAG

Research · Model or response · Open · English legal · 2024

6,858 expert-annotated query-answer pairs with character-level evidence spans.

Reported results

No comparable score table is recorded in this collection.

What the benchmark measures

6,858 expert-annotated query-answer pairs with character-level evidence spans.

The evaluated unit is a model response or component output. Read the source for the exact prompt, tool and harness conditions.

How it is scored

Character-level precision and recall.

Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.

Sources

Tests retrieval and grounding at character level.