Anthropic Evaluations

This skill should be used when the user asks to "create evals", "evaluate an agent", "build evaluation suite", or mentions agent testing, graders, or benchmarks. Also suggest when building coding agents, conversational agents, or research agents that need quality assurance. Use when this capability is needed.

tomevault-io a8c029c 2 files · 4.3 KB Updated

File contents

tomevault-io/skills-registry/tree/main/dwmkerr--claude-toolkit--anthropic-evaluations commit a8c029c1e2

Frequently asked questions

npx skillmds@latest add tomevault-io/anthropic-evaluations