Model Compare

Blind side-by-side multi-model comparison. Send one prompt to 2-4 models simultaneously, present responses anonymously (Model A / B / C / D), let the user pick a winner, then reveal identities and show which model won. Supports custom evaluation criteria, synthesis of responses, and vote history logging. Trigger when the user says "compare models", "test these models", "which model is better for", "A/B test", "blind comparison", "model evaluation", or wants to see how different AI models handle the same prompt. Can also be used for prompt engineering — testing how different models interpret the same instructions. Supports reasoning-effort A/B via provider:@effort:model_id specs.

moonlight-lupin 258ca3f 12 files · 199.4 KB Updated

File contents

moonlight-lupin/agent-skills/tree/main/mlops/model-compare commit 258ca3f410

Frequently asked questions

npx skillmds@latest add moonlight-lupin/model-compare