Mobile Gui Agent Eval

Evaluates the capability of multimodal large language models to act as autonomous mobile GUI agents. It probes their ability to plan tasks, predict correct interaction types, accurately ground UI elements, and successfully complete complex, multi-step workflows in both static and dynamic Android environments. Use when the user wants to benchmark on AndroidControl, AndroidLab, Android Agent Arena (A3), or asks about evaluating this task. Reports Success Rate (SR).

qhjqhj00 f3079ef 4.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mobile-gui-agent-eval commit f3079ef3ee

Frequently asked questions

npx skillmds add qhjqhj00/mobile-gui-agent-eval