Alrm Manipulation Eval

Evaluates an agentic LLM's ability to plan and execute multistep robotic manipulation tasks in simulation. It tests closed-loop reasoning via ReAct-style loops, comparing code-as-policy and tool-as-policy execution modes across linguistically diverse tasks. Use when the user wants to benchmark on ALRM Simulation Benchmark, or asks about evaluating this task. Reports task_completion.

qhjqhj00 84c9d7c 2.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/alrm-manipulation-eval commit 84c9d7cb88

Frequently asked questions

npx skillmds add qhjqhj00/alrm-manipulation-eval