# Oc Agent Evaluation

> Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics.

- Skill: `luokai0/oc-agent-evaluation` (Agent Skill)
- Install (CLI): `npx skillmds@latest add luokai0/oc-agent-evaluation`
- Raw SKILL.md: https://api.skillmd.com/api/skills/luokai0/oc-agent-evaluation/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: DevOps & Infra
- Author: luokai0 (https://skillmd.com/u/luokai0)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/luokai0/oc-agent-evaluation

---


# Agent Evaluation

## Overview
Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics.

## Before Starting
1. What specific task do you need this skill for?
2. What inputs are available?
3. What is the expected output format?

## Usage
Install via: `npx clawhub install agent-evaluation`
Documentation: https://clawskills.sh/skills/agent-evaluation

## Core Functionality
Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics.

## Best Practices
- Test in sandbox environment first
- Check latest version before production use
- Review source docs for advanced configuration

## Related Skills
- agent-installer
- clawhub-publisher

