Marioqa Eval

Evaluates a model's ability to perform video question answering with varying levels of temporal reasoning complexity. It probes whether models can correctly link visual events in gameplay videos to answer questions that require single-frame, event-level, or multi-step causal/temporal understanding. Use when the user wants to benchmark on MarioQA, or asks about evaluating this task. Reports accuracy.

qhjqhj00 feec83e 2.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/marioqa-eval commit feec83ee21

Frequently asked questions

npx skillmds add qhjqhj00/marioqa-eval