Autored Eval

Evaluates the safety alignment and vulnerability of large language models against adversarial red-teaming prompts. It measures how effectively generated or human-crafted harmful instructions can bypass safety filters to elicit unsafe model responses. Use when the user wants to benchmark on AutoRed & Baseline Red-Teaming Datasets, or asks about evaluating this task. Reports Attack Success Rate (ASR).

qhjqhj00 b69cb9e 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/autored-eval commit b69cb9ed97

Frequently asked questions

npx skillmds add qhjqhj00/autored-eval