Tagalong Dojo Eval

Evaluates the effectiveness of an adversarial agent in jailbreaking safety-aligned operator agents through conversational interaction. It measures how well a small attacker model can trigger prohibited tool usage on unseen malicious tasks using reinforcement learning. Use when the user wants to benchmark on TagAlong-Dojo, or asks about evaluating this task. Reports Attack Success Rate (ASR).

qhjqhj00 b79bf74 4.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/tagalong-dojo-eval commit b79bf74f4c

Frequently asked questions

npx skillmds add qhjqhj00/tagalong-dojo-eval