Agentsynth Eval

Evaluates the ability of multimodal language models to execute long-horizon, multi-step computer-use tasks on a desktop environment. It probes visual grounding, precise GUI interaction, state tracking, and error recovery across varying task complexities and software domains. Use when the user wants to benchmark on AgentSynth, or asks about evaluating this task. Reports success rate.

qhjqhj00 bc874ef 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/agentsynth-eval commit bc874ef64e

Frequently asked questions

npx skillmds add qhjqhj00/agentsynth-eval