# Arxiv Gamedevbench Evaluating Agentic Capabili

> Learned from arXiv paper GameDevBench: Evaluating Agentic Capabilities Through Game Development. Use this skill to scaffold Node.js experiments based on the paper method.

- Skill: `johnalbertini14-glitch/arxiv-gamedevbench-evaluating-agentic-capabili` (Agent Skill, multi-file: 6 files)
- Install (CLI): `npx skillmds@latest add johnalbertini14-glitch/arxiv-gamedevbench-evaluating-agentic-capabili`
- Raw SKILL.md: https://api.skillmd.com/api/skills/johnalbertini14-glitch/arxiv-gamedevbench-evaluating-agentic-capabili/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: johnalbertini14-glitch (https://skillmd.com/u/johnalbertini14-glitch)
- Updated: 2026-09-09
- Page: https://skillmd.com/skills/johnalbertini14-glitch/arxiv-gamedevbench-evaluating-agentic-capabili

---


# arxiv-gamedevbench-evaluating-agentic-capabili

## Source
- Paper key: 44f3ad505bee7a5c25a60d2a3686cb7e
- Title: GameDevBench: Evaluating Agentic Capabilities Through Game Development
- Categories: cs.AI,cs.CL,cs.SE

## Learned insight
Despite rapid progress on coding agents, progress on their multimodal counterparts has lagged behind. A key challenge is the scarcity of evaluation testbeds that combine the complexity of software development with the need for deep multimodal understanding. Game development provides such a testbed as agents must navigate large, dense codebases while manipulating intrinsically multimodal assets such as shaders, sprites, and animations within a visual game scene. We present GameDevBench, the first

## Node.js implementation entry
`node {baseDir}/scripts/run.js`

