Legal AI Solution Design MapInterpreting benchmarks for legal work and system design.
Menu

BFCL v4

Agent · Agent or completed task · Open · Domain-general · 2024

Tool-calling tasks assessed by checking the selected functions and their arguments.

Reported results

What the benchmark measures

Tool-calling tasks assessed by checking the selected functions and their arguments.

The evaluated unit is an agent or completed task. Read the source for the exact prompt, tool and harness conditions.

How it is scored

Function correctness and argument accuracy rates.

Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.

Sources

Standard function-calling benchmark.