Benchmark Models

Cross-model benchmark for vibestack skills. Runs the same prompt through Claude, GPT (via Codex CLI), and Gemini side-by-side — compares latency, tokens, cost, and optionally quality via LLM judge. Answers "which model is actually best for this skill?" with data instead of vibes.

timurgaleev 436c264 7.6 KB Updated

File contents

timurgaleev/vibestack/tree/main/skills/benchmark-models commit 436c26405b

Frequently asked questions

npx skillmds@latest add timurgaleev/benchmark-models