AI Forge Eval

Behavioral eval for skills and agents — whether the artifact actually changes model behavior, not just whether it scores well on a rubric. Use when verifying a skill or agent works in practice, or when tracking its performance across repeated runs to catch regressions. Triggers are test this skill, test this agent, eval this, does this skill work, behavioral eval, run eval, benchmark, benchmark this skill, track regression, verify outputs, does this agent work. Don't use for rubric-only scoring — that's ai-forge-judge.

robcsaszar c9291ec 8 files · 43.5 KB Updated

File contents

robcsaszar/ai-forge/tree/main/skills/ai-forge-eval commit c9291ec28f

Frequently asked questions

npx skillmds@latest add robcsaszar/ai-forge-eval