MCP Mark Comprehensive Agent Benchmark

Evaluate LLM agents through realistic multi-turn tool-use workflows across 127 complex MCP tasks spanning CRUD operations, state management, and error handling. Use when assessing agent capabilities on real-world tool orchestration beyond shallow read-only interactions.

adu2021 994c4c1 3.5 KB Updated

File contents

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.2-claude-opus-4.6/mcp-mark-comprehensive-agent-benchmark commit 994c4c1a6d

Frequently asked questions

npx skillmds@latest add adu2021/mcp-mark-comprehensive-agent-benchmark