Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

LegalLens

Reasoning · Model or response · Open · US consumer · 2024

Tests identification of legal violations in text through named entity recognition (NER) and natural language inference (NLI).

Reported results

What the benchmark measures

Tests identification of legal violations in text through named entity recognition (NER) and natural language inference (NLI).

The evaluated unit is a model response or component output. Read the source for the exact prompt, tool and harness conditions.

How it is scored

Weighted F1 for NER, macro F1 for NLI.

Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.

Sources

NLLP Workshop 2024 shared task.