Xstest Eval

Probes whether large language models exhibit exaggerated safety behaviors by refusing safe prompts due to lexical overfitting or system prompt effects. It measures the model's ability to distinguish between genuinely unsafe requests and safe prompts that merely resemble unsafe content. Use when the user wants to benchmark on XSTest, or asks about evaluating this task. Reports response_classification.

qhjqhj00 026be4d 5.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/xstest-eval commit 026be4d8c8

Frequently asked questions

npx skillmds add qhjqhj00/xstest-eval