Llava Bench Eval

Assesses multimodal chatbot capabilities, including conversation, detailed description, and complex visual reasoning. It measures how well a model follows instructions and understands novel or challenging visual inputs compared to a strong text-only baseline. Use when the user wants to benchmark on LLaVA-Bench, or asks about evaluating this task. Reports relative_score.

qhjqhj00 d91ad1c 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/llava-bench-eval commit d91ad1c86d

Frequently asked questions

npx skillmds add qhjqhj00/llava-bench-eval