Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

About Liz Chase

I’m a lawyer and senior legal engineer working on how legal work translates into AI systems and how those systems should be tested. I built this map because benchmark scores are useful, but they only tell us what was measured. My focus is on connecting those results to the work lawyers actually do: the skills involved, what remains untested, and what those gaps mean when designing and evaluating legal AI systems. Benchmark results on this site come from their cited sources; the skills, task mappings, gaps and design responses are my own analysis.

Resources

  • LegalBench — the open-source benchmark for legal reasoning, with 162 tasks contributed by legal professionals and researchers. GitHub repository
  • Inspect AI — the UK AI Security Institute's open-source framework for building and running model and agent evaluations, including tool use and multi-turn tasks. GitHub repository
  • AgentBench — research benchmark for evaluating models acting as agents across environments including databases, knowledge graphs, operating systems and web tasks. GitHub repository
  • METR Task Standard — a practical specification for defining reproducible tasks and environments for agent evaluation. GitHub repository
  • OpenAI Evals — an open-source framework and benchmark registry for building task-specific evaluations of models and LLM systems. GitHub repository

Corrections, better evidence and useful research are always welcome.

Connect with me on LinkedIn