Mobile Bench Eval

Evaluates LLM-based mobile agents on real-world task execution across single-app and multi-app scenarios. It probes the agent's ability to plan, navigate UIs, call APIs, and collaborate across applications to complete user-defined goals. Use when the user wants to benchmark on Mobile-Bench, or asks about evaluating this task. Reports PassRate.

qhjqhj00 84c758e 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mobile-bench-eval commit 84c758e9bd

Frequently asked questions

npx skillmds add qhjqhj00/mobile-bench-eval