Widget Captioning Eval

This benchmark evaluates a model's ability to generate natural language descriptions for individual mobile UI elements using multimodal inputs. It probes the capability to fuse visual appearance and structural hierarchy data to produce accurate, context-aware captions for accessibility and UI understanding tasks. Use when the user wants to benchmark on Widget Captioning Dataset, or asks about evaluating this task. Reports CIDEr.

qhjqhj00 e5a3b02 4.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/widget-captioning-eval commit e5a3b02f28

Frequently asked questions

npx skillmds add qhjqhj00/widget-captioning-eval