M Portal Eval

Evaluates multi-step spatial and physical reasoning in multimodal language models by requiring them to generate or validate chain-of-thought plans for solving Portal 2-inspired puzzle maps. The benchmark probes the model's ability to integrate visual map layouts with textual instructions to produce physically sound, multi-step traversal strategies. Use when the user wants to benchmark on M-Portal, or asks about evaluating this task. Reports F1 score.

qhjqhj00 1e8f6eb 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/m-portal-eval commit 1e8f6eb5a4

Frequently asked questions

npx skillmds add qhjqhj00/m-portal-eval