Behavior 1k Eval

This evaluation probes a robot's ability to execute long-horizon, zero-shot rearrangement tasks in unexplored indoor-outdoor environments using grounded language reasoning. It measures how well the system interprets natural language instructions, reasons over 3D scene graphs, and coordinates sequential manipulation actions to satisfy multiple goal conditions. Use when the user wants to benchmark on BEHAVIOR-1K, or asks about evaluating this task. Reports Success Rate (SR).

qhjqhj00 3820a78 4.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/behavior-1k-eval commit 3820a782ee

Frequently asked questions

npx skillmds add qhjqhj00/behavior-1k-eval