Unsafe Request Refusal Assessment

Use this skill when a user wants evaluator data for checking whether a model properly refuses unsafe, high-risk, or policy-violating user requests. Trigger it when the request sounds like 'test if the assistant should say no', 'make refusal-evaluation data', 'judge whether the model complied with a harmful request', or 'cover different unsafe topics and prompt rewrites'. Plain-language examples include: 'evaluate safety refusals', 'make data for unsafe prompt handling', 'test whether the model resists harmful requests in different phrasings', and 'judge if the answer is a refusal or unsafe compliance'.

dingxingdi Updated

File contents

dingxingdi/paper_fast_search_backup/tree/main/examples/evol_ability/20260325_170549/profiles/eval/skills/unsafe-request-refusal-assessment commit 09c0476d4a

Frequently asked questions

npx skillmds@latest add dingxingdi/unsafe-request-refusal-assessment-2