Chinesafe Eval

This benchmark evaluates the safety of large language models in Chinese by testing their ability to correctly classify text as safe or unsafe across multiple sensitive categories. It probes whether models can reliably detect harmful, policy-violating, or sensitive content in a Chinese-language context using both generation-based and perplexity-based strategies. Use when the user wants to benchmark on ChineseSafe, or asks about evaluating this task. Reports accuracy.

qhjqhj00 91143dc 4.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/chinesafe-eval commit 91143dc4e2

Frequently asked questions

npx skillmds add qhjqhj00/chinesafe-eval