Multi Label Toxicity Detection Eval

Evaluates an LLM's ability to identify multiple concurrent toxicity categories in real-world prompts using a fine-grained 15-category taxonomy. It probes fine-grained safety alignment, multi-label classification under ambiguous annotations, and the model's robustness to sparse or noisy supervision signals. Use when the user wants to benchmark on Q-A-MLL, H-X-MLL, R-A-MLL, or asks about evaluating this task. Reports mean Average Precision.

qhjqhj00 630ad45 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/multi-label-toxicity-detection-eval commit 630ad4567a

Frequently asked questions

npx skillmds add qhjqhj00/multi-label-toxicity-detection-eval