Eval Driven Dev

Instrument Python LLM apps, build golden datasets, write eval-based tests, run them, and root-cause failures — covering the full eval-driven development cycle. Make sure to use this skill whenever a user is developing, testing, QA-ing, evaluating, or benchmarking a Python project that calls an LLM, even if they don't say "evals" explicitly. Use for making sure an AI app works correctly, catching regressions after prompt changes, debugging why an agent started behaving differently, or validating output quality before shipping.

dvcrn Updated 32 repo stars

File contents

dvcrn/openclaw-skills-marketplace/tree/main/plugins/yiouli--eval-driven-dev/skills/eval-driven-dev commit ee5e024f8b

Frequently asked questions

npx skillmds@latest add dvcrn/eval-driven-dev