Findingdory Eval

This benchmark evaluates long-term memory and spatio-temporal reasoning in embodied agents. It requires agents to recall specific past interactions from a video history to select goal frames and navigate to target entities in dynamic, photorealistic environments over long-horizon tasks. Use when the user wants to benchmark on FindingDory, or asks about evaluating this task. Reports LL-SR.

qhjqhj00 df7b23e 4.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/findingdory-eval commit df7b23e12d

Frequently asked questions

npx skillmds add qhjqhj00/findingdory-eval