Prompt Injection Review Eval

This benchmark evaluates the vulnerability of large language models to prompt injection attacks when generating scientific paper reviews. It probes whether hidden or biased instructions embedded in parsed PDFs can systematically skew the model's review scores and recommendations. Use when the user wants to benchmark on ICLR 2024 Review Dataset, or asks about evaluating this task. Reports Rating.

qhjqhj00 2c0c375 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/prompt-injection-review-eval commit 2c0c3759f4

Frequently asked questions

npx skillmds add qhjqhj00/prompt-injection-review-eval