Livemcpbench Eval

Evaluates LLM agents' capability to dynamically discover, select, and chain Model Context Protocol (MCP) tools to complete complex, multi-step real-world tasks. It probes meta-tool learning and multi-tool collaboration in large-scale, time-varying tool ecosystems. Use when the user wants to benchmark on LiveMCPBench, or asks about evaluating this task. Reports task success rate.

qhjqhj00 dfc47f2 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/livemcpbench-eval commit dfc47f2c13

Frequently asked questions

npx skillmds add qhjqhj00/livemcpbench-eval