Omniear Eval

Evaluates embodied agent reasoning by testing how well models infer capability gaps, dynamic tool acquisition, and coordination needs from environmental constraints. Probes the ability to ground abstract reasoning in physical reality under partial observability. Use when the user wants to benchmark on EAR-Bench, or asks about evaluating this task. Reports Success Rate (SR).

qhjqhj00 965f8e9 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/omniear-eval commit 965f8e9adf

Frequently asked questions

npx skillmds add qhjqhj00/omniear-eval