Model Evaluation Runner

This skill runs model evaluation suites to measure accuracy, fairness, and performance metrics. Use when asked to evaluate a model, benchmark against baselines, or assess model fairness. Also consider when a model is promoted without evaluation evidence. Suggest when the user trains a model without defining evaluation criteria.

mittuled 88c5e3b 6 files · 24.8 KB Updated

File contents

mittuled/skill-os/tree/main/agents/engineering/ai-ml-engineer/model-evaluation-runner commit 88c5e3b515

Frequently asked questions

npx skillmds@latest add mittuled/model-evaluation-runner