Agentharm Eval

This benchmark evaluates the harmfulness and safety alignment of LLM-based agents by measuring their compliance with malicious, multi-step tasks that require coherent tool chaining. It probes whether models can be coerced into executing harmful behaviors through direct prompting or simple jailbreak templates, while tracking refusal rates and capability preservation. Use when the user wants to benchmark on AgentHarm, or asks about evaluating this task. Reports harm score.

qhjqhj00 45f9e6a 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/agentharm-eval commit 45f9e6a427

Frequently asked questions

npx skillmds add qhjqhj00/agentharm-eval