manipulationnet-eval
ManipulationNet: An Infrastructure for Benchmarking Real-World Robot Manipulation with Physical Skill Challenges and Embodied Multimodal Reasoning — Chen et al. (2026) (arXiv:2603.04363, 2026)
What this evaluates
Evaluates real-world robot manipulation capabilities across two complementary tracks: physical skills (sensorimotor execution under contact, clearance, and perceptual constraints) and embodied reasoning (multimodal grounding of natural language and visual instructions into grounded actions).
Datasets
- ManipulationNet Benchmark — total ?; splits: test (-1)
Metrics
declutter rate— range: [0, 1]- Ratio of objects successfully grasped and removed to the total number of spawned objects in a scene.
grasp success rate— range: [0, 1]- Ratio of successful grasps to the total number of grasp attempts made during a scene.
time efficiency— range: seconds- Total wall-clock time required to complete the task or scene layout.
task success rate(primary) — range: [0, 1]- Binary or graded success determined by centralized judges applying task-specific metrics after integrity verification.
Input / output format
Input: Standardized physical object set configured on a workstation; task instructions (language prompts, visual prompts, or projected scene layouts) delivered via mnet-client; external camera recording setup.
Output: Robot execution logs, real-time status messages, cryptographic hashes of video frames, and a complete recorded video of the trial.
Scoring recipe
def score_submission(video, logs, task_goal):
if not verify_one_time_code(video): return None
if not verify_hashes(video, logs): return None
if task_goal == 'grasping_in_clutter':
declutter_rate = count_success(logs) / count_spawned(logs)
grasp_success_rate = count_success(logs) / count_attempts(logs)
time_eff = logs['end_time'] - logs['start_time']
return declutter_rate, grasp_success_rate, time_eff
else:
return judge_task_success(video, logs, task_goal)
Common pitfalls
- Submissions must display a server-generated one-time code in the camera view; missing it invalidates the trial.
- Integrity relies on real-time cryptographic hashes of frames and a final video hash; pre-recording or post-processing is explicitly blocked.
- Metrics are not auto-computed by the client; final scoring requires centralized committee verification against task-specific criteria.
Evidence (verbatim from paper)
The grasping in clutter benchmark evaluates the grasping performance based on the declutter rate (the ratio of objects successfully grasped and removed to all spawned objects), the grasp success rate (the ratio of successful grasps to the total number of grasp attempts), and time efficiency. Task-specific performance metrics are applied only after these integrity checks have been passed, ensuring objective and comparable evaluations across systems.
Citation
@misc{chen2026manipulationnet,
title={ManipulationNet: An Infrastructure for Benchmarking Real-World Robot Manipulation with Physical Skill Challenges and Embodied Multimodal Reasoning},
author={Chen et al. (2026)},
year={2026},
note={arXiv:2603.04363}
}
- arXiv: 2603.04363