Ragen Agent Eval

Evaluates LLM agents' multi-turn decision-making and reasoning capabilities across symbolic planning, risk-sensitive reasoning, and realistic web interaction environments. It probes the agent's ability to complete interactive tasks under noisy or probabilistic feedback while maintaining exploration and training stability. Use when the user wants to benchmark on Bandit, Sokoban, Frozen Lake, WebShop, or asks about evaluating this task. Reports success rate.

qhjqhj00 a2a68d3 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/ragen-agent-eval commit a2a68d37e7

Frequently asked questions

npx skillmds add qhjqhj00/ragen-agent-eval