Duobench

Benchmark planner×implementer LLM pairings on a real GitHub issue and chart quality-per-dollar. Use when the user says things like "benchmark <models> on duobench", "run a duobench eval", "add <planner>/<implementer> to the eval", "produce plots about the duobench results", or "re-plot the last duobench run". You orchestrate plan→implement→judge phase jobs in tmux, aggregate results.json, then write seaborn plots from results.json/trial.json.

alejandro-ao a161566 27 files · 164.0 KB Updated

File contents

alejandro-ao/duobench/tree/main/skills/duobench commit a161566bfd

Frequently asked questions

npx skillmds@latest add alejandro-ao/duobench