Phishnchips Eval

Evaluates the security and robustness of autonomous LLM email agents against phishing attacks by measuring how different system prompt configurations affect detection sensitivity and operational false positive rates. It specifically probes the model's ability to maintain high recall while minimizing usability costs, and tests adversarial brittleness under infrastructure phishing conditions where attacker-controlled domains match sender addresses. Use when the user wants to benchmark on Synthetic Email Phishing Corpus, or asks about evaluating this task. Reports Net Effectiveness (Recall-FPR).

qhjqhj00 540022b 4.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/phishnchips-eval commit 540022baf0

Frequently asked questions

npx skillmds add qhjqhj00/phishnchips-eval