Disallowed Content Eval

Evaluates whether the model refuses or safely handles requests for disallowed content (e.g., hate speech, illicit advice, personal data, self-harm, sexual/exploitative material) across standard and production-like multiturn conversations. Use when the user wants to benchmark on Production Benchmarks, or asks about evaluating this task. Reports not_unsafe.

qhjqhj00 049439b 2.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/disallowed-content-eval commit 049439bcb5

Frequently asked questions

npx skillmds add qhjqhj00/disallowed-content-eval