Odysseyarena Benchmarking Long Horizon Active

Design and run inductive agent benchmarks where LLMs must discover hidden rules through long-horizon interaction loops rather than following explicit instructions. Use when the user mentions 'inductive agent evaluation', 'long-horizon benchmarking', 'hidden rule discovery', 'active exploration benchmark', 'OdysseyArena', or 'agent world-model induction'.

ndpvt-web Updated

File contents

ndpvt-web/arxiv-claude-skills/tree/main/skills/odysseyarena-benchmarking-long-horizon-active commit a94a2e594d

Frequently asked questions

npx skillmds@latest add ndpvt-web/odysseyarena-benchmarking-long-horizon-active