Agentic Evaluation

Implements systematic evaluation, benchmarking, and testing of AI agent behaviors with tool-use accuracy, hallucination detection, multi-turn reasoning metrics, and automated grading pipelines.

paulpas ec4f090 31.3 KB Updated

File contents

paulpas/agent-skill-router/tree/main/skills/agent/agentic-evaluation commit ec4f090414

Frequently asked questions

npx skillmds@latest add paulpas/agentic-evaluation