# Evaluate model-generated code execution with SandboxFusion

> Use SandboxFusion to run and judge LLM-generated code in controlled sandboxes across many languages and benchmark-style evaluation tasks.

- Skill: `agentskillexchange/evaluate-model-generated-code-execution-with-sandboxfusion` (Agent Skill)
- Install (CLI): `npx skillmds@latest add agentskillexchange/evaluate-model-generated-code-execution-with-sandboxfusion`
- Raw SKILL.md: https://api.skillmd.com/api/skills/agentskillexchange/evaluate-model-generated-code-execution-with-sandboxfusion/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- Author: agentskillexchange (https://skillmd.com/u/agentskillexchange)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/agentskillexchange/evaluate-model-generated-code-execution-with-sandboxfusion

---


# Evaluate model-generated code execution with SandboxFusion

Use SandboxFusion to run and judge LLM-generated code in controlled sandboxes across many languages and benchmark-style evaluation tasks.

## Prerequisites

Docker, or conda and Poetry for manual installation

## Installation

Use the upstream install or setup path that matches your environment:
- docker build -f ./scripts/Dockerfile.base -t code_sandbox:base .
- docker build -f ./scripts/Dockerfile.server -t code_sandbox:server .
- docker run -d --rm -p 8080:8080 code_sandbox:server make run-online
- conda create -n sandbox -y python=3.12

Requirements and caveats from upstream:
- Python (python, pytest)
- Python (GPU)
- **Online Judge**: Implementation of Evaluation & RL datasets that requires code running

Basic usage or getting-started notes:
- **Code Runner**: Run and return the result of a code snippet
- Build the image locally:
- sed -i '1s/.*/FROM code_sandbox:base/' ./scripts/Dockerfile.server

- Source: https://github.com/bytedance/SandboxFusion
- Extracted from upstream docs: https://raw.githubusercontent.com/bytedance/SandboxFusion/HEAD/README.md

## Documentation

- https://bytedance.github.io/SandboxFusion/

## Source

- [Agent Skill Exchange](https://agentskillexchange.com/skills/evaluate-model-generated-code-execution-with-sandboxfusion/)

