Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

HELM Enterprise (Legal)

Reasoning · Model or response · Open · Primarily US · 2024

IBM extension of Stanford HELM with legal-specific scenarios.

Reported results

No comparable score table is recorded in this collection.

What the benchmark measures

IBM extension of Stanford HELM with legal-specific scenarios.

The evaluated unit is a model response or component output. Read the source for the exact prompt, tool and harness conditions.

How it is scored

HELM 7-metric framework.

Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.

Sources

Assesses legal scenarios using several measures of performance.