Beep Korean Toxic Speech Eval

This benchmark evaluates models' ability to detect social bias (gender and other types) and hate speech in Korean online news comments. It probes whether models can distinguish between hate speech, offensive language, and neutral comments, and whether incorporating bias labels improves hate speech detection. Use when the user wants to benchmark on BEEP!, or asks about evaluating this task. Reports F1.

qhjqhj00 fe22515 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/beep-korean-toxic-speech-eval commit fe225156c5

Frequently asked questions

npx skillmds add qhjqhj00/beep-korean-toxic-speech-eval