Visual Text Grounding Eval

Evaluates multimodal large language models' ability to perform precise spatial reasoning and visual text grounding in document images. It tests whether models can generate accurate bounding boxes that support their textual answers, both from scratch (OCR-free) and when provided with OCR text (OCR-based), while also measuring their instruction-following capability. Use when the user wants to benchmark on ChartQA, DocVQA, InfographicsVQA, TRINS, or asks about evaluating this task. Reports IoU.

qhjqhj00 af39cc7 3.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/visual-text-grounding-eval commit af39cc7c1a

Frequently asked questions

npx skillmds add qhjqhj00/visual-text-grounding-eval