Eval Driven Dev

Instrument Python LLM apps, build golden datasets, write eval-based tests, run them, and root-cause failures — covering the full eval-driven development cycle. Make sure to use this skill whenever a user is developing, testing, QA-ing, evaluating, or benchmarking a Python project that calls an LLM, even if they don't say "evals" explicitly. Use for making sure an AI app works correctly, catching regressions after prompt changes, debugging why an agent started behaving differently, or validating output quality before shipping.

williamlimasilva e3bc7f0 30 files · 241.4 KB Updated

File contents

williamlimasilva/.copilot/tree/main/skills/eval-driven-dev commit e3bc7f0a94

Frequently asked questions

npx skillmds add williamlimasilva/eval-driven-dev