Molmospaces Bench Eval

Evaluates zero-shot generalization of vision-language-action and navigation policies across diverse indoor scenes. Probes robustness to environmental perturbations, language prompt variations, and sim-to-real transferability for long-horizon manipulation and semantic navigation tasks. Use when the user wants to benchmark on MolmoSpaces-Bench, or asks about evaluating this task. Reports success rate.

qhjqhj00 f69a61c 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/molmospaces-bench-eval commit f69a61c141

Frequently asked questions

npx skillmds add qhjqhj00/molmospaces-bench-eval