Backdoor Detection Purification Eval

Evaluates language models' vulnerability to backdoor attacks and the effectiveness of detection and purification defenses. It probes whether a model can correctly classify clean text while resisting trigger-induced misclassifications, and whether a defense can identify poisoned samples without degrading benign task performance. Use when the user wants to benchmark on SST-2, YELP, AG’s News, or asks about evaluating this task. Reports AUC.

qhjqhj00 1d87703 4.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/backdoor-detection-purification-eval commit 1d87703558

Frequently asked questions

npx skillmds add qhjqhj00/backdoor-detection-purification-eval