Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

LegalAgentBench

Agent · Agent or completed task · Open · Chinese legal · 2025

37 tools for interacting with legal knowledge bases.

Reported results

What the benchmark measures

37 tools for interacting with legal knowledge bases.

The evaluated unit is an agent or completed task. Read the source for the exact prompt, tool and harness conditions.

How it is scored

Task completion rate across tool-use scenarios.

Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.

Sources

Published at ACL 2025; tests completion of legal tasks using tools.