Adversarial Ml Robustness

Model-level adversarial machine learning and robustness — at the bar of a researcher who breaks and defends models for a living. Use when threat-modeling an ML/LLM model, evaluating or claiming robustness, or defending against evasion / adversarial examples (FGSM, PGD, C&W, transferable & patch/ physical attacks), data poisoning, backdoors/trojans (clean-label, trigger-based), model extraction/ stealing, model inversion, membership inference, or LLM jailbreaks / training-data extraction. Covers the NIST Adversarial ML taxonomy (AI 100-2e2025) and MITRE ATLAS; threat models (white/black/gray-box, Lp budgets); defenses and their limits (adversarial training, certified/randomized smoothing, why defensive distillation & gradient masking are false security); and the core skill — honest robustness evaluation with adaptive attacks, AutoAttack, and RobustBench. Defensive, not an attack playbook.

sanjeevrg89 6c47e8b 5 files · 53.4 KB Updated

File contents

sanjeevrg89/arete/tree/main/skills/adversarial-ml-robustness commit 6c47e8b4d0

Frequently asked questions

npx skillmds@latest add sanjeevrg89/adversarial-ml-robustness