LePhantomCite
1,300 legal brief excerpts containing deliberately introduced citation problems.
Reported results
Agentic citation-verification F1 · higher is better
| Claude Code, Opus 4.8P 76.1 · R 62.8 | 68.8 | |
|---|---|---|
| GPT-5P 40.8 · R 84.4 | 55 | |
| Qwen3.6-27BP 28.9 · R 65.0 | 40 | |
| GPT-OSS-120BP 21.1 · R 55.1 | 30.5 | |
| Gemini 2.5 FlashP 16.9 · R 66.9 | 27 | |
| Qwen3-8BP 12.0 · R 41.1 | 18.6 |
F1 balances precision and recall when identifying citation problems. The paper also studies generated citations over time. Those results are separate from this verification table.
What the benchmark measures
1,300 legal brief excerpts containing deliberately introduced citation problems.
The evaluated unit is a model response or component output. Read the source for the exact prompt, tool and harness conditions.
How it is scored
Precision and recall by hallucination category.
Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.
Sources
- Primary benchmark source
- LePhantomCite paper, Table 2, Jun 2026 · results 2026-06 · checked 2026-09-18
Tests whether systems can identify fabricated or otherwise problematic citations.