Comparative Response Ranking

Use this skill when a user wants evaluator data for ranking two or more open-ended answers by overall quality, especially when people say things like 'rank these outputs', 'choose the better answer', 'sort multiple responses from best to worst', or 'compare several candidates, not just one'. Trigger it for pairwise or listwise judging where the evaluator must order responses by helpfulness, relevance, completeness, or usability. Plain-language examples include: 'make answer-ranking data', 'test which of three responses is best', 'evaluate multiple candidates at once', and 'create hard pairwise judge examples where both answers look okay'.

dingxingdi a7dc4eb 5 files · 14.5 KB Updated

File contents

dingxingdi/paper_fast_search_backup/tree/main/examples/evol_ability/20260325_170549/profiles/eval/skills/comparative-response-ranking commit a7dc4eb542

Frequently asked questions

npx skillmds@latest add dingxingdi/comparative-response-ranking-2