Agent Evaluation

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks

welitonevoc 798fdf0 35.7 KB Updated 1 repo stars

File contents

welitonevoc/Biblia-Codex/tree/main/agente-IA/skills/agent-evaluation commit 798fdf0968

Frequently asked questions

npx skillmds add welitonevoc/agent-evaluation