Team 1-800-Shared-Tasks (fine-tuned Phi-3 Medium) · Results from 2024-10
LegalLens
Tests identification of legal violations in text through named entity recognition (NER) and natural language inference (NLI).
Reported results
What the benchmark measures
Tests identification of legal violations in text through named entity recognition (NER) and natural language inference (NLI).
The evaluated unit is a model response or component output. Read the source for the exact prompt, tool and harness conditions.
How it is scored
Weighted F1 for NER, macro F1 for NLI.
Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.