Automated Red Teaming Eval

Evaluates an LLM's capability to generate effective adversarial prompts (red teaming attacks) for arbitrary safety goals. It measures both the success rate of eliciting targeted behaviors and the diversity of the generated attacks across in-domain and out-of-domain objectives. Use when the user wants to benchmark on garak adversarial goals, or asks about evaluating this task. Reports attack success rate.

qhjqhj00 21745a6 4.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/automated-red-teaming-eval commit 21745a6839

Frequently asked questions

npx skillmds add qhjqhj00/automated-red-teaming-eval