Mobileagentbench Eval

This benchmark evaluates the performance of LLM-based mobile agents on Android GUI navigation tasks. It measures end-to-end task completion, action efficiency, response latency, computational cost, and the agent's ability to correctly determine task completion without stopping too early or too late. Use when the user wants to benchmark on MobileAgentBench, or asks about evaluating this task. Reports Success Rate (SR).

qhjqhj00 a92f171 4.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mobileagentbench-eval commit a92f1719d2

Frequently asked questions

npx skillmds add qhjqhj00/mobileagentbench-eval