Skill Benchmarking

Run skill benchmarks with discriminating-only assertions against evals.json for any model and any AI agent. Use when benchmarking a skill against a model not yet tested, running with_skill/without_skill eval pairs, producing benchmark-<model>.json, re-grading an existing run, adding Phase 2 model comparison results, reviewing results in the eval viewer, updating README benchmark tables, or cleaning non-discriminating assertions from evals.json. Enforces strict grader isolation (the context that generates responses never grades them) and evidence-only passing (assertions pass only on explicit content, never on implication or charity). Works with Claude Code, Gemini CLI, GitHub Copilot, Cursor, and any other AI coding assistant.

rusel95 60a2862 16 files · 172.7 KB Updated

File contents

rusel95/ios-agent-skills/tree/main/scripts/benchmarking commit 60a2862c59

Frequently asked questions

npx skillmds@latest add rusel95/skill-benchmarking