Shawshank Bench Eval

Evaluates the vulnerability of embodied AI agents (Vision-Language Models) to indirect environmental jailbreaks, where malicious instructions are physically embedded in the environment (e.g., on walls or tables) rather than provided as direct text prompts. It measures both the agent's susceptibility to harmful behavior and the collateral impact on benign task execution. Use when the user wants to benchmark on Shawshank-Bench, or asks about evaluating this task. Reports ASR.

qhjqhj00 2f58eb7 4.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/shawshank-bench-eval commit 2f58eb7f90

Frequently asked questions

npx skillmds add qhjqhj00/shawshank-bench-eval