Gui Ceval Eval

Evaluates multimodal large language models and agents on Chinese mobile GUI interaction tasks. It probes atomic capabilities like visual perception, grounding, and planning, as well as end-to-end execution reliability in both offline simulation and real-device online environments. Use when the user wants to benchmark on GUI-CEval, or asks about evaluating this task. Reports Online Agent success rate.

qhjqhj00 9f4dbb7 4.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/gui-ceval-eval commit 9f4dbb7201

Frequently asked questions

npx skillmds add qhjqhj00/gui-ceval-eval