Bfcl Eval

Evaluates an LLM's ability to generate correct function calls from natural language prompts, covering single, multiple, parallel, and parallel-multiple API invocations across different programming languages. Use when the user wants to benchmark on Berkeley Function-Calling Benchmark (BFCL), or asks about evaluating this task. Reports Overall Accuracy.

qhjqhj00 0e2b08d 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/bfcl-eval commit 0e2b08d9d2

Frequently asked questions

npx skillmds add qhjqhj00/bfcl-eval