Titi Jailbreak Eval

Evaluates LLM safety alignment against stateless multi-turn adversarial attacks. It measures how often models generate unsafe responses and which specific risk categories they fail on when subjected to iterative, context-independent prompt injection. Use when the user wants to benchmark on ModifiedMasterKeyJailbreakQuestions, or asks about evaluating this task. Reports Unsafe Response Rate.

qhjqhj00 3c91feb 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/titi-jailbreak-eval commit 3c91feb027

Frequently asked questions

npx skillmds add qhjqhj00/titi-jailbreak-eval