Beat Backdoor Detection Eval

Evaluates the ability of a black-box defense mechanism to detect backdoor-unaligned samples in LLMs by measuring changes in the model's refusal behavior when a malicious probe is concatenated to the input. Use when the user wants to benchmark on MaliciousInstruct + Advbench + UltraChat-200k, or asks about evaluating this task. Reports AUROC.

qhjqhj00 c899a01 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/beat-backdoor-detection-eval commit c899a01bb5

Frequently asked questions

npx skillmds add qhjqhj00/beat-backdoor-detection-eval