Benchmark Models

Cross-model benchmark for gstack skills. Runs the same prompt through Claude, GPT (via Codex CLI), and Gemini side-by-side — compares latency, tokens, cost, and optionally quality via LLM judge. Answers "which model is actually best for this skill?" with data instead of vibes. Separate from /benchmark, which measures web page performance. Use when: "benchmark models", "compare models", "which model is best for X", "cross-model comparison", "model shootout". (gstack)

kitfunso 11d0eeb 31.4 KB Updated

File contents

kitfunso/claude-config/tree/main/skills/benchmark-models commit 11d0eeba50

Frequently asked questions

npx skillmds@latest add kitfunso/benchmark-models