Marvel Eval

Evaluates multimodal large language models on multidimensional abstract visual reasoning and perceptual grounding. It probes the model's ability to recognize complex geometric and abstract patterns, track temporal/spatial changes, and perform multi-step visual reasoning across diverse puzzle configurations. Use when the user wants to benchmark on MARVEL, or asks about evaluating this task. Reports accuracy.

qhjqhj00 904096e 4.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/marvel-eval commit 904096eae0

Frequently asked questions

npx skillmds add qhjqhj00/marvel-eval