Tower Mind Agent Benchmark

Evaluate LLM agent capabilities using tower defense game environment with multimodal observations (pixel, text, structured state). Benchmark reveals critical agent limitations: inadequate planning validation, inflexible decision-making, and inefficient action use. Demonstrates significant performance gap between current LLMs and human experts, providing structured framework for measuring agent planning, adaptation, and hallucination tendencies.

adu2021 6a7a4c5 6.9 KB Updated

File contents

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.2-claude-opus-4.6/tower-mind-agent-benchmark commit 6a7a4c5e06

Frequently asked questions

npx skillmds@latest add adu2021/tower-mind-agent-benchmark