Suede AI Eval

Suede Labs AI eval design and coverage audit: AI-SPEC, failure-mode rubric with severity scoring, concrete pass/fail eval cases, coverage and infrastructure scores, and mechanical acceptance gates. Use when a change ships LLM, RAG, agent, classifier, prompt, or generated-media behavior, or when asked to write evals for an AI feature, design test cases for a model surface, audit existing eval coverage, or judge whether AI behavior is safe to ship. No AI-SPEC means no eval plan, and no eval plan holds the recommended ship verdict. NOT FOR: reviewing or grading the implementation code behind the AI surface (use suede-code); wiring a passing suite into CI as a required check (use suede-ci-gate); UAT of the built feature beyond the eval suite (a private Suede Labs companion, not in this pack).

tuyv 19bcbf0 4 files · 40.5 KB Updated

File contents

tuyv/ccpm/tree/main/preset-registry/skills/jasoncolapietro-suede-creator-skills-suede-ai-eval commit 19bcbf0813

Frequently asked questions

npx skillmds@latest add tuyv/suede-ai-eval