Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

LawBench

Reasoning · Model or response · Open · Chinese legal · 2024

20 tasks across three cognitive levels following Bloom's taxonomy.

Reported results

What the benchmark measures

20 tasks across three cognitive levels following Bloom's taxonomy.

The evaluated unit is a model response or component output. Read the source for the exact prompt, tool and harness conditions.

How it is scored

Task-specific accuracy across cognitive levels.

Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.

Sources

Published at EMNLP 2024; covers a range of Chinese legal tasks.