Deception Eval

Probes the model's tendency to exhibit deceptive alignment, including alignment faking in chain-of-thought reasoning, jailbreak success rates, and strategic behavior shifts between evaluation and deployment stages. Use when the user wants to benchmark on DECEPTIONBENCH, StrongReject, JailbreakBench, BeaverTails, HarmfulQA, or asks about evaluating this task. Reports DTR.

qhjqhj00 beb221e 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/deception-eval commit beb221e8c5

Frequently asked questions

npx skillmds add qhjqhj00/deception-eval