Benchmark And Baseline Selector

Given a research proposal, claim, method, or result, recommend the baselines and benchmarks needed to support it — split into a MINIMAL set (necessary for credibility, skipping any risks rejection) and a SUGGESTED set (complementary, strengthens external validity). Use whenever the user proposes something and asks "what should I compare against", "what baselines do I need", "which benchmarks support this claim", "is this enough to be convincing", "what's the minimal experiment set", "what would a reviewer demand", or shares a method/idea and wants the comparison plan. Default to running this on any proposal so claims are backed by the right evidence. Always check first whether the input is sufficient (precise claim, field, the novel component, metric); if not, request the missing context before recommending rather than guessing. Use experiment-design to design the run itself (or to design a measurement when no benchmark exists), and hypothesis-and-ablation-planner for internal ablations.

jurgendn Updated

File contents

jurgendn/agent-skills/tree/main/skills/research-experimentation/benchmark-and-baseline-selector commit 46075baf30

Frequently asked questions

npx skillmds@latest add jurgendn/benchmark-and-baseline-selector