Mulberry Eval

Evaluates the multimodal reasoning and understanding capabilities of MLLMs across diverse domains including mathematics, chart interpretation, scientific/medical images, and hallucination detection. It measures how well models generate step-by-step reasoning paths and reflect on errors to produce correct answers. Use when the user wants to benchmark on MathVista, MMStar, MMMU, ChartQA, DynaMath, HallBench, MM-Math, MMEsum, or asks about evaluating this task. Reports Average Benchmark Score.

qhjqhj00 f4ebc27 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mulberry-eval commit f4ebc27342

Frequently asked questions

npx skillmds add qhjqhj00/mulberry-eval