Swarmbench Eval

This benchmark evaluates emergent decentralized coordination in LLM-driven multi-agent systems under strict local perception and communication constraints. It simulates five canonical swarm tasks—Pursuit, Synchronization, Foraging, Flocking, and Transport—in a 2D grid environment to probe whether LLMs can form adaptive group strategies and execute robust long-range planning without global information. Use when the user wants to benchmark on SwarmBench, or asks about evaluating this task. Reports Performance.

qhjqhj00 eacdfb6 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/swarmbench-eval commit eacdfb6ec5

Frequently asked questions

npx skillmds add qhjqhj00/swarmbench-eval