Gameplayqa Eval

GameplayQA evaluates multi-modal large language models' ability to understand decision-dense, first-person synchronized multi-video environments. It probes capabilities in agent-state tracking, temporal reasoning, and cross-video event alignment across three cognitive difficulty levels. Use when the user wants to benchmark on GameplayQA, or asks about evaluating this task. Reports accuracy.

qhjqhj00 185f2f6 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/gameplayqa-eval commit 185f2f6c17

Frequently asked questions

npx skillmds add qhjqhj00/gameplayqa-eval