Adversarial Defense Eval

Evaluates the robustness, seamlessness, and general utility of LLMs against adversarial inputs (jailbreaks, toxicity, hallucinations, bias) using an inference-time defense framework. Use when the user has predictions and gold and needs to compute robustness score.

qhjqhj00 b65a231 4.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/adversarial-defense-eval commit b65a23105a

Frequently asked questions

npx skillmds add qhjqhj00/adversarial-defense-eval