Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

LePhantomCite

Assurance · Model or response · Mixed · US litigation · 2026

1,300 legal brief excerpts containing deliberately introduced citation problems.

Reported results

Agentic citation-verification F1 · higher is betterResults 2026-06 · checked 2026-09-18
LePhantomCite, citation verification
Claude Code, Opus 4.8P 76.1 · R 62.8
68.8
GPT-5P 40.8 · R 84.4
55
Qwen3.6-27BP 28.9 · R 65.0
40
GPT-OSS-120BP 21.1 · R 55.1
30.5
Gemini 2.5 FlashP 16.9 · R 66.9
27
Qwen3-8BP 12.0 · R 41.1
18.6

F1 balances precision and recall when identifying citation problems. The paper also studies generated citations over time. Those results are separate from this verification table.

Six agentic systems evaluated on 1,300 legal brief excerpts with injected citation problems. Source: LePhantomCite paper, Table 2, Jun 2026.

What the benchmark measures

1,300 legal brief excerpts containing deliberately introduced citation problems.

The evaluated unit is a model response or component output. Read the source for the exact prompt, tool and harness conditions.

How it is scored

Precision and recall by hallucination category.

Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.

Sources

Tests whether systems can identify fabricated or otherwise problematic citations.