Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

PLawBench

Reasoning · Model or response · Open · Chinese legal · 2026

Assesses language models on legal practice tasks using detailed scoring criteria.

Reported results

Top recorded result69.7 Overall rubric score, weighted across three tasks

GPT-5.2 · Results from 2026-01

PLawBench paper, Table 3

What the benchmark measures

Assesses language models on legal practice tasks using detailed scoring criteria.

The evaluated unit is a model response or component output. Read the source for the exact prompt, tool and harness conditions.

How it is scored

Rubric-guided scores across practice dimensions.

Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.

Sources

Uses assessment rubrics intended to reflect how lawyers evaluate work.