Mandela

Audit an evaluation, benchmark, or scoring harness for leakage, answering whether outside ground truth actually enters or the system is grading itself. Use when building or reviewing an eval or benchmark, or when the user says "is this eval leaking" or "check this benchmark for contamination". It reports where ground truth enters and where it does not, and it edits nothing.

majiayu000 35331a5 2 files · 3.3 KB Updated 567 repo stars

File contents

majiayu000/claude-skill-registry-data/tree/main/design/mandela commit 35331a5643

Frequently asked questions

npx skillmds add majiayu000/mandela