Visualwebbench Eval

Evaluates multimodal LLMs' ability to understand web pages and ground UI elements. It probes capabilities across seven subtasks including image captioning, web question answering, OCR, element/action grounding, and action prediction. Use when the user wants to benchmark on VisualWebBench, or asks about evaluating this task. Reports Average Score.

qhjqhj00 abf4054 2.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/visualwebbench-eval commit abf40545ec

Frequently asked questions

npx skillmds add qhjqhj00/visualwebbench-eval