Evaluate

Score a candidate on a split with honest, variance-aware evaluation. Use whenever you need a number for a candidate (the algorithm calls it internally; you can also call it directly to inspect). Runs the target via the adapter for each task, scores each rollout, aggregates mean + standard error, and reports pass^k when trials > 1. Never touches the test split (that is finalize's sealed job).

skillberry-ai 07ea0cd 7 files · 27.7 KB Updated

File contents

skillberry-ai/cap-evolve/tree/main/skills/phases/evaluate commit 07ea0cde74

Frequently asked questions

npx skillmds@latest add skillberry-ai/evaluate