Model Bench

Run a 3-leg evaluation against a newly released LLM. Leg 1 is trap questions with deterministic answers, leg 2 is an agentic orchestration task with a planted failure, leg 3 is a brownfield feature build in a real repo. Use when the user wants to vibe-check, test, evaluate, or compare a new frontier model, or mentions model-bench, model vibe check, or launch-day protocol.

nateslabach 9a00284 7 files · 13.2 KB Updated

File contents

nateslabach/skills/tree/main/skills/model-bench commit 9a002844dc

Frequently asked questions

npx skillmds@latest add nateslabach/model-bench