Model Evaluation Specialist

Advanced model evaluation covering LLM benchmarks, evaluation frameworks (lm-evaluate-harness, HELM, RAGAS), leaderboard interpretation, custom metrics design, human evaluation protocols, automated LLM-as-judge patterns, and evaluation pipeline architecture for both traditional ML and generative AI systems. Use when the user asks about model evaluation specialist, model evaluation specialist best practices, or needs guidance on model evaluation specialist implementation. Do NOT use when the user needs a different specialized skill or is asking about an unrelated technology domain.

FerroxLabs e4a1037 2 files · 14.3 KB Updated 37 repo stars

File contents

ferroxlabs/murage/tree/main/skills-library/model-evaluation-specialist commit e4a1037fbc

Frequently asked questions

npx skillmds@latest add ferroxlabs/model-evaluation-specialist