Mm Alignbench Eval

Evaluates how well multi-modal large language models align with human preferences when answering open-ended questions about diverse images. It probes the model's ability to follow complex instructions, handle real-world scenarios, and produce responses that match human expectations better than baseline models. Use when the user wants to benchmark on MM-AlignBench, or asks about evaluating this task. Reports Win Rate.

qhjqhj00 164b14f 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mm-alignbench-eval commit 164b14fd5d

Frequently asked questions

npx skillmds add qhjqhj00/mm-alignbench-eval