Quality Evals

Use when a team wants to validate test-runner's bug-detection accuracy or test-author's authoring fidelity against their OWN app before trusting the manual-qa bundle on real work — "how good is this agent, really", "benchmark test-author/test-runner", "build a gold suite", "self-eval the manual QA team". Ports a held-out-answer-key eval methodology (deterministic Tier A scoring + a judged Tier B rubric) that keeps a self-authored eval honest, generalized to any app.

arozumenko Updated

File contents

arozumenko/sdlc-skills/tree/main/bundles/manual-qa/skills/quality-evals commit bba12063c0

Frequently asked questions

npx skillmds@latest add arozumenko/quality-evals