Multi Turn Preference Evaluation

Use this skill when a user wants data where the evaluator must judge which assistant handled a back-and-forth conversation better, especially when people say things like 'compare two chats', 'see who handled the follow-up better', 'test memory across turns', 'check whether the answer stayed on track', or 'make the second question matter'. Trigger it for pairwise judging of multi-turn dialogues where later turns depend on earlier turns, and where the decision should reflect instruction following, coherence, recall, and usefulness across the full exchange. Example triggers in plain language include: 'give me judge data for two-step conversations', 'make questions where the follow-up exposes weak memory', 'compare which answer handles the second part better', and 'test whether the model keeps the context straight across turns'.

dingxingdi dbbbe8c 5 files · 16.4 KB Updated

File contents

dingxingdi/paper_fast_search_backup/tree/main/examples/evol_ability/20260325_170549/profiles/eval/skills/multi-turn-preference-evaluation commit dbbbe8c09d

Frequently asked questions

npx skillmds@latest add dingxingdi/multi-turn-preference-evaluation-2