Gui Agent Halluc Eval

This protocol evaluates GUI agents on visual grounding, action execution, and hallucination rates across mobile, desktop, and web interfaces. It measures how well models localize UI elements, execute multi-step tasks under varying instruction granularities, and avoid perception or reasoning errors. Use when the user wants to benchmark on ScreenSpot-V2, ScreenSpot-Pro, AndroidControl, GUI-Odyssey, or asks about evaluating this task. Reports Action Type Accuracy (Type), Grounding Accuracy (GR), Step-wise Success Rate (SR), Hallucination Rate (HR).

qhjqhj00 73a1f32 3.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/gui-agent-halluc-eval commit 73a1f325f5

Frequently asked questions

npx skillmds add qhjqhj00/gui-agent-halluc-eval