Fara 7b Agentic Eval

This evaluation probes the agentic capabilities of computer-use models by measuring their ability to complete multi-step web browsing and task-completion tasks on live websites. It assesses both functional success rates and operational efficiency, including token usage, cost, and interaction length. Use when the user wants to benchmark on WebVoyager, Online-Mind2Web, DeepShop, WebTailBench, ScreenSpot, or asks about evaluating this task. Reports success rate.

qhjqhj00 0523216 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/fara-7b-agentic-eval commit 0523216d32

Frequently asked questions

npx skillmds add qhjqhj00/fara-7b-agentic-eval