Calvin Long Horizon Eval

Evaluates a robot policy's ability to chain multiple sub-goals sequentially using only visual observations and goal images. It probes long-horizon planning, goal-conditioned control, and the capacity to generalize to unseen goal configurations without explicit reward signals. Use when the user wants to benchmark on CALVIN, or asks about evaluating this task. Reports Success rate.

qhjqhj00 e367d8b 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/calvin-long-horizon-eval commit e367d8b3d8

Frequently asked questions

npx skillmds add qhjqhj00/calvin-long-horizon-eval