LLM Eval Suite Scaffolder

Stand up an evaluation suite for an LLM feature from scratch — a representative dataset, the right metrics, a baseline score, and a CI gate — using DeepEval, promptfoo, or RAGAS. Use when a feature has no evals, before tuning a prompt, or when adding an LLM feature to CI.

imtiazrayhan Updated

File contents

imtiazrayhan/agentscamp-library/tree/main/skills/llm-eval-suite-scaffolder commit 738f55a2e7

Frequently asked questions

npx skillmds@latest add imtiazrayhan/llm-eval-suite-scaffolder