Mas Bench Eval

Evaluates the ability of mobile GUI agents to complete complex, real-world automation tasks across single-app and cross-app scenarios. It specifically probes how well agents can integrate predefined or self-generated shortcuts (APIs, deep links, RPA scripts) with standard GUI interactions to improve task success, execution efficiency, and cost-effectiveness. Use when the user wants to benchmark on MAS-Bench, or asks about evaluating this task. Reports SR.

qhjqhj00 feabd14 4.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mas-bench-eval commit feabd14a33

Frequently asked questions

npx skillmds add qhjqhj00/mas-bench-eval