Agent Evaluation

Run one specified Test Agent on one specified Benchmark Case exactly once, privately score that execution, and return one protocol result.

prism-shadow b2356a3 9.3 KB Updated

File contents

prism-shadow/penguin-harness/tree/main/plugins/agent-tuning/skills/agent-evaluation commit b2356a3052

Frequently asked questions

npx skillmds@latest add prism-shadow/agent-evaluation