# Evaluate long-horizon agents against WildClawBench

> Use WildClawBench to benchmark agents on hard end-to-end OpenClaw tasks covering tool orchestration, multimodal work, coding, safety, and long-horizon planning.

- Skill: `agentskillexchange/evaluate-long-horizon-agents-against-wildclawbench` (Agent Skill)
- Install (CLI): `npx skillmds@latest add agentskillexchange/evaluate-long-horizon-agents-against-wildclawbench`
- Raw SKILL.md: https://api.skillmd.com/api/skills/agentskillexchange/evaluate-long-horizon-agents-against-wildclawbench/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: agentskillexchange (https://skillmd.com/u/agentskillexchange)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/agentskillexchange/evaluate-long-horizon-agents-against-wildclawbench

---


# Evaluate long-horizon agents against WildClawBench

Use WildClawBench to benchmark agents on hard end-to-end OpenClaw tasks covering tool orchestration, multimodal work, coding, safety, and long-horizon planning.

## Prerequisites

WildClawBench assets; OpenClaw environment; target agent/model under test

## Installation

No source-backed install or usage instructions could be extracted automatically. Review the upstream project before running this skill in a sensitive workflow.

- Source: https://github.com/InternLM/WildClawBench

## Documentation

- https://internlm.github.io/WildClawBench/

## Source

- [Agent Skill Exchange](https://agentskillexchange.com/skills/evaluate-long-horizon-agents-against-wildclawbench/)

