Mandela

Audit any eval, metric, experiment, or benchmark for leakage — does external ground-truth enter independently, or are the model, scorer, and designer just confirming a result no outside truth ever produced? Use before trusting any 'how we'll know it worked' — an A/B, a holdout, a score, a validation — and whenever a result feels too clean or self-confirming. Walks an 8-pattern leakage taxonomy and returns only the patterns that fire, each with an independence fix. Read-only.

lilmgenius 59e9a5c 3.2 KB Updated

File contents

lilmgenius/paperthin/tree/main/skills/depth/mandela commit 59e9a5c8d3

Frequently asked questions

npx skillmds@latest add lilmgenius/mandela