Cultureguard Multilingual Safety Eval

Evaluates multilingual content safety guard models on their ability to detect harmful or unsafe prompts and responses across diverse languages and cultural contexts, including zero-shot generalization to unseen languages. Use when the user wants to benchmark on CultureGuard, PolyGuardPrompts, RTP-LX, MultiJail, XSafety, Aya Red-teaming, or asks about evaluating this task. Reports harmful-F1.

qhjqhj00 ae2991b 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/cultureguard-multilingual-safety-eval commit ae2991b49b

Frequently asked questions

npx skillmds add qhjqhj00/cultureguard-multilingual-safety-eval