Mme Cot Eval

Evaluates the quality, robustness, and efficiency of Chain-of-Thought reasoning in Large Multimodal Models. It probes whether models can generate accurate intermediate reasoning steps, maintain performance consistency between direct and CoT prompting, and produce relevant, non-redundant reasoning traces. Use when the user wants to benchmark on MME-CoT, or asks about evaluating this task. Reports F1 score.

qhjqhj00 5a57dae 4.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mme-cot-eval commit 5a57dae34a

Frequently asked questions

npx skillmds add qhjqhj00/mme-cot-eval