Steer Bench Eval

Evaluates the safety and helpfulness alignment of multimodal large language models under single-turn versus multi-turn interactive settings, specifically probing the static-to-dynamic generalization gap and the evolution of safety failure rates across conversation turns. Use when the user wants to benchmark on Steer-Bench, or asks about evaluating this task. Reports pass_rate.

qhjqhj00 b0f1a95 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/steer-bench-eval commit b0f1a95416

Frequently asked questions

npx skillmds add qhjqhj00/steer-bench-eval