Ml Benchmark Evaluation

Rigorous methodology for evaluating ML models on established benchmarks. Covers proper train/val/test splits, baseline verification from original papers, exact metric formula discrepancies, data-leak detection checklist, multi-seed robustness, and honest reporting templates. Use when claiming to beat published baselines, writing methods papers, or auditing existing results.

synthetic-sciences 32e777c 9.0 KB Updated

File contents

synthetic-sciences/openscience/tree/main/backend/cli/skills/ml-training/ml-benchmark-evaluation commit 32e777cc51

Frequently asked questions

npx skillmds@latest add synthetic-sciences/ml-benchmark-evaluation