Mobile R1 Benchmark Eval

Evaluates a vision-language model's ability to navigate and complete multi-step tasks on mobile GUIs. It measures step-level action correctness, full trajectory success, and robustness to intermediate errors in a simulated Android environment. Use when the user wants to benchmark on Chinese Mobile Agent Benchmark, or asks about evaluating this task. Reports Accuracy (Acc.).

qhjqhj00 cad95fc 3.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mobile-r1-benchmark-eval commit cad95fc271

Frequently asked questions

npx skillmds add qhjqhj00/mobile-r1-benchmark-eval