Model Evaluation

Build or revise an evidence-linked capability profile for a model execution setup by running repeated, representative real tasks under matched conditions. Use when comparing models, providers, plans, coding harnesses, or prompt/tool profiles for task allocation; when asking "which model is good enough for this work?", "evaluate this model", "compare model capability", "模型评测/能力画像/模型适合什么任务", or whether a characterized setup may have degraded. Do not use for provider setup, public leaderboard summaries, one-off response review, automatic routing, or a degradation verdict without an accepted baseline.

lidessen 40fb66a 4 files · 22.8 KB Updated

File contents

lidessen/rossovia/tree/main/skills/model-evaluation commit 40fb66a5ec

Frequently asked questions

npx skillmds@latest add lidessen/model-evaluation