Evaluate long-horizon agents against WildClawBench

Use WildClawBench to benchmark agents on hard end-to-end OpenClaw tasks covering tool orchestration, multimodal work, coding, safety, and long-horizon planning.

agentskillexchange Updated 28 repo stars

File contents

Evaluate long-horizon agents against WildClawBench

Use WildClawBench to benchmark agents on hard end-to-end OpenClaw tasks covering tool orchestration, multimodal work, coding, safety, and long-horizon planning.

Prerequisites

WildClawBench assets; OpenClaw environment; target agent/model under test

Installation

No source-backed install or usage instructions could be extracted automatically. Review the upstream project before running this skill in a sensitive workflow.

Documentation

Source

agentskillexchange/skills/tree/main/skills/evaluate-long-horizon-agents-against-wildclawbench commit 460ab872dc

Frequently asked questions

npx skillmds@latest add agentskillexchange/evaluate-long-horizon-agents-against-wildclawbench