Agent Evaluation And Benchmarking

Design and interpret evaluations of AI coding-agent behavior across realistic tasks, correctness, safety, usefulness, and maintainability. Use when creating benchmarks, comparing agents or skills, measuring regressions, or making claims about agent capability.

ChloeVPin 8c5e077 6.1 KB Updated

File contents

ChloeVPin/open-agent-skills/tree/main/skills/agent-evaluation-and-benchmarking commit 8c5e077bd7

Frequently asked questions

npx skillmds@latest add chloevpin/agent-evaluation-and-benchmarking