Mirrorbench Eval

This benchmark evaluates self-centric intelligence and mirror self-recognition in Multimodal Large Language Models (MLLMs) within an embodied simulation. It probes the model's ability to perform self-referential reasoning and navigate tasks under varying cognitive difficulty levels and body configurations (humanoid vs. robotic). Use when the user wants to benchmark on MirrorBench, or asks about evaluating this task. Reports AVG.

qhjqhj00 fd0abfc 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mirrorbench-eval commit fd0abfc113

Frequently asked questions

npx skillmds add qhjqhj00/mirrorbench-eval