Orak Eval

Evaluates LLM agents' long-horizon decision-making and gameplay capabilities across 12 diverse video games spanning six genres. It probes the effectiveness of different agentic strategies (zero-shot, reflection, planning, skill-management) and the impact of multimodal inputs (text vs. visual) on action inference. The benchmark also assesses generalization to unseen in-game scenarios, out-of-distribution games, and non-game tasks like math and web navigation. Use when the user wants to benchmark on Orak, or asks about evaluating this task. Reports normalization score.

qhjqhj00 29ab0b5 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/orak-eval commit 29ab0b5e06

Frequently asked questions

npx skillmds add qhjqhj00/orak-eval