Vision Language Response Judgment

Use this skill when a user wants evaluator data for judging answers about images, charts, infographics, screenshots, diagrams, or other visual inputs. Trigger it when people use plain requests like 'compare who answered the picture question better', 'score the chart explanation', 'rank several image-based answers', or 'make judge data for multimodal responses'. It is especially appropriate for scoring evaluation, pair comparison, or batch ranking in vision-language tasks where the judge must reason over both the visual content and the textual response. Example triggers include: 'evaluate answers about charts', 'judge which caption is better', 'rank image question responses', and 'make harder visual judge items with hallucinations'.

dingxingdi ae89923 5 files · 14.5 KB Updated

File contents

dingxingdi/paper_fast_search_backup/tree/main/examples/evol_ability/20260325_170549/profiles/eval/skills/vision-language-response-judgment commit ae89923288

Frequently asked questions

npx skillmds@latest add dingxingdi/vision-language-response-judgment-2