Detecting LLM Peer Reviews Eval

Evaluates the efficacy of covert watermarking techniques embedded in manuscript PDFs to force LLM-generated peer reviews to contain specific hidden markers. It also tests the robustness of these watermarks against common reviewer defenses like paraphrasing, detection prompts, and page cropping, as well as the performance of cryptic prompt injection via gradient-based optimization. Use when the user wants to benchmark on ICLR 2024 submissions, ICLR 2021 submissions, ICLR 2024 submissions (control), NSF Grant Proposals, PRC 2022 abstracts, PeerRead papers, or asks about evaluating this task. Reports Accuracy.

qhjqhj00 5ae16ad 5.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/detecting-llm-peer-reviews-eval commit 5ae16ad8d7

Frequently asked questions

npx skillmds add qhjqhj00/detecting-llm-peer-reviews-eval