Eval Authoring

This skill should be used when the user mentions "llm eval", "evaluation", "promptfoo", "deepeval", "regression test", "llm judge", "golden dataset", "eval suite", "test a prompt", or wants to prove a prompt/model change improved rather than regressed behavior. It provides a standardized methodology for authoring assertion-based, LLM-as-judge, and golden-dataset eval suites that gate CI.

sigistry d684a46 3 files · 15.6 KB Updated

File contents

sigistry/marketplace/tree/main/plugins/llm-app-hardener/skills/eval-authoring commit d684a46bc0

Frequently asked questions

npx skillmds@latest add sigistry/eval-authoring