Compositeharm Eval

Evaluates cross-lingual safety degradation in LLMs by measuring how well models refuse harmful prompts and avoid generating unsafe content across English and five Indic languages. Use when the user wants to benchmark on CompositeHarm, or asks about evaluating this task. Reports Refusal Rate (RR), Attack Success Rate (ASR).

qhjqhj00 8d4fb82 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/compositeharm-eval commit 8d4fb82aab

Frequently asked questions

npx skillmds add qhjqhj00/compositeharm-eval