Thyme Multimodal Eval

Evaluates multimodal large language models on image manipulation, visual perception, mathematical reasoning, and general vision-language tasks. It probes whether autonomous code generation and execution for image processing improves downstream accuracy and reduces hallucination. Use when the user wants to benchmark on MME-RealWorld, HR Bench, MathVista, Hallucination bench, MMStar, or asks about evaluating this task. Reports accuracy.

qhjqhj00 9cd1501 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/thyme-multimodal-eval commit 9cd150103a

Frequently asked questions

npx skillmds add qhjqhj00/thyme-multimodal-eval