Adversarial Text Attack Eval

Evaluates the robustness of BERT-based text classifiers against word-level adversarial attacks by measuring how well perturbed inputs maintain semantic meaning and syntactic structure while successfully flipping model predictions. It compares three attack methods across three standard classification benchmarks to determine the optimal balance between attack success, semantic preservation, and computational efficiency. Use when the user wants to benchmark on IMDB, AG News, SST2, or asks about evaluating this task. Reports Attack Accuracy.

qhjqhj00 efd56ec 4.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/adversarial-text-attack-eval commit efd56ecb84

Frequently asked questions

npx skillmds add qhjqhj00/adversarial-text-attack-eval