Deepeyes Eval

Evaluates large vision-language models on fine-grained visual perception, grounding, hallucination mitigation, and multimodal reasoning. It specifically probes the model's ability to autonomously use image zoom-in tools for interleaved visual-linguistic reasoning (iMCoT) to solve high-resolution and complex visual tasks. Use when the user wants to benchmark on V* Bench, HR-Bench, refCOCO / refCOCO+ / refCOCOg / ReasonSeg, POPE, MathVista, MathVerse, or asks about evaluating this task. Reports accuracy.

qhjqhj00 fbdc711 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/deepeyes-eval commit fbdc71103e

Frequently asked questions

npx skillmds add qhjqhj00/deepeyes-eval