Run repeatable model and agent eval suites and inspect scoring traces with Inspect AI

Run benchmark-style eval suites against models or agents, then inspect scored traces instead of relying on ad hoc chats and gut feel.

agentskillexchange Updated 28 repo stars

File contents

agentskillexchange/skills/tree/main/skills/run-repeatable-model-and-agent-eval-suites-and-inspect-scoring-traces-with-inspect-ai commit 6aab7ad98d

Frequently asked questions

npx skillmds@latest add agentskillexchange/run-repeatable-model-and-agent-eval-suites-and-inspect-scori