Eval Driven Dev

Add instrumentation, build golden datasets, write eval-based tests, run them, root-cause failures, and iterate — Ensure your Python LLM application works correctly. Make sure to use this skill whenever a user is developing, testing, QA-ing, evaluating, or benchmarking a Python project that calls an LLM. Use for making sure an LLM application works correctly, catching regressions after prompt changes, fixing unexpected behavior, or validating output quality before shipping.

mit-network Updated 2 repo stars

File contents

mit-network/awesome-copilot-skills/tree/main/skills/eval-driven-dev commit 51e0253963

Frequently asked questions

npx skillmds@latest add mit-network/eval-driven-dev