Eval

Evaluate and rank agent results by metric or LLM judge for an AgentHub session. Use when the user runs /hub:eval or asks to score, compare, or pick a winner among completed AgentHub agents.

sinhoneyy Updated 11 repo stars

File contents

sinhoneyy/master-skills/tree/master/skills/eval commit 1eaf3c5dd8

Frequently asked questions

npx skillmds@latest add sinhoneyy/eval