Agent Evaluation

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks

thedixitjain 72b300c 36.7 KB Updated 2 repo stars

File contents

thedixitjain/the-mega-skill-library/tree/main/library/ai-agents-and-harness/agent-evaluation commit 72b300cf36

Frequently asked questions

npx skillmds add thedixitjain/agent-evaluation