Mousi Vlm Eval

Evaluates multimodal understanding and reasoning across visual question answering, OCR, region-level VQA, and visual conversation tasks. It probes how effectively poly-visual expert ensembles fuse information from multiple encoders compared to single-expert baselines. Use when the user wants to benchmark on LLaVA-1.5 Benchmark Suite, or asks about evaluating this task. Reports accuracy.

qhjqhj00 280b7a8 2.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mousi-vlm-eval commit 280b7a89e7

Frequently asked questions

npx skillmds add qhjqhj00/mousi-vlm-eval