Prompt Attack Detection Eval

This benchmark evaluates an LLM's or monitoring system's ability to distinguish between safe user inputs and malicious prompt injection attacks. It measures both false positive rates on legitimate interactions and false negative rates on adversarial prompts to assess overall security robustness. Use when the user wants to benchmark on Gandalf, Tensor-Trust, SPML-Dataset, or asks about evaluating this task. Reports Error Rate (ER).

qhjqhj00 0341e81 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/prompt-attack-detection-eval commit 0341e81ca0

Frequently asked questions

npx skillmds add qhjqhj00/prompt-attack-detection-eval