Benchmarks
25
legal and adjacent
Model evaluation
19
scored on the answer
Agent evaluation
6
scored on completed work
Selected skills
0
profile builder

Agent benchmarks judge completed work product, not a single answer, so the ceilings are far lower. No frontier model finishes even 10% of Harvey LAB end to end.

Scored on

BenchmarkLayerSkills and taskScoringAccessNotes

Two axes. Layer is the subject a benchmark covers; scored on is how it marks the work. They are independent: RedlineBench and Legal Research Bench score agent execution, but their subjects are Contract and Research, which is why 6 benchmarks are agent-scored while only 4 sit in the Agent layer.

Entries curated from primary sources: peer-reviewed papers, official repositories, and maintainer documentation. Task descriptions, capability mappings and layer grouping are editorial. Data as at September 2026.