Claude Opus 4.5, native function calling · Results from 2026-04
BFCL v4
Tool-calling tasks assessed by checking the selected functions and their arguments.
Reported results
What the benchmark measures
Tool-calling tasks assessed by checking the selected functions and their arguments.
The evaluated unit is an agent or completed task. Read the source for the exact prompt, tool and harness conditions.
How it is scored
Function correctness and argument accuracy rates.
Scores remain in the original unit. They are not normalised or combined with results from other benchmarks.