Agent Evaluation

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks

gabrielmoreira Updated 17 repo stars

File contents

gabrielmoreira/agent-skills-mirror/tree/main/mirrors/repos/fabioc-aloha@Alex_Skill_Mall/plugins/ai-agents/agent-evaluation/skills/agent-evaluation commit e22147f037

Frequently asked questions

npx skillmds@latest add gabrielmoreira/agent-evaluation