Uinnexus Eval

Evaluates mobile agents' ability to complete long-horizon, dependency-rich tasks on real mobile applications. It specifically probes atomic-to-compositional generalization, testing how well agents handle task concatenation, context transitions, and deep analysis across different app types and languages. Use when the user wants to benchmark on UI-NEXUS, or asks about evaluating this task. Reports Success Rate.

qhjqhj00 0d0a60b 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/uinnexus-eval commit 0d0a60b0c8

Frequently asked questions

npx skillmds add qhjqhj00/uinnexus-eval