Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

30 benchmarks · sources checked 22 September 2026

Aced the exam, but
can legal AI do the job?

Benchmaxxed describes AI that is optimised to perform well on benchmarks rather than necessarily becoming better at real-world work. This site maps 30 legal AI benchmark results to actual legal skills and tasks to test what those scores really tell us about the current state of legal AI, and offer design suggestions for known trouble spots.

One evaluation, two denominators

94.5%of individual criteria passed
same system, same evaluation
25.4%of whole tasks resolved

Muse Spark 1.2 on Harvey’s Legal Agent Benchmark, VALS AI, reported 2026-09-15, checked 2026-09-19. A task counts only if every criterion passes. Passing most checks still leaves the job unfinished. Read the scoring method.

Benchmark coverage by task

ScoredMapped, no score—Nothing recorded