Agent Red Teaming Eval

Evaluates the security and robustness of frontier AI agents against adversarial prompt injections and policy violations across multiple models and realistic deployment scenarios. It measures how effectively attacks transfer between models and whether model capability or inference compute correlates with safety. Use when the user wants to benchmark on Agent Red Teaming (ART) benchmark, or asks about evaluating this task. Reports ASR.

qhjqhj00 0b051e9 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/agent-red-teaming-eval commit 0b051e92fc

Frequently asked questions

npx skillmds add qhjqhj00/agent-red-teaming-eval