Agent benchmarks judge completed work product, not a single answer, so the ceilings are far lower. No frontier model finishes even 10% of Harvey LAB end to end.
Scored on
| Benchmark | Layer | Skills and task | Scoring | Access | Notes |
|---|
Agent benchmarks judge completed work product, not a single answer, so the ceilings are far lower. No frontier model finishes even 10% of Harvey LAB end to end.
| Benchmark | Layer | Skills and task | Scoring | Access | Notes |
|---|
Two axes. Layer is the subject a benchmark covers; scored on is how it marks the work. They are independent: RedlineBench and Legal Research Bench score agent execution, but their subjects are Contract and Research, which is why 6 benchmarks are agent-scored while only 4 sit in the Agent layer.
Entries curated from primary sources: peer-reviewed papers, official repositories, and maintainer documentation. Task descriptions, capability mappings and layer grouping are editorial. Data as at September 2026.