Dialogue Safety Robustness Eval

Evaluates the robustness of offensive language detection models against adversarial human attacks in single-turn and multi-turn dialogue contexts. It measures classifier resilience when exposed to iterative, context-aware attacks designed to evade safety filters. Use when the user wants to benchmark on Wikipedia Toxic Comments, or asks about evaluating this task. Reports Weighted-F1.

qhjqhj00 dc107bb 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/dialogue-safety-robustness-eval commit dc107bb8a6

Frequently asked questions

npx skillmds add qhjqhj00/dialogue-safety-robustness-eval