Toxicity Detection Eval

Probes the ability of text generation models to produce non-toxic content by measuring average toxicity scores and comparing them via statistical significance testing. It specifically evaluates how accounting for classifier uncertainty affects the reliability of these comparisons. Use when the user wants to benchmark on BOLD, RealToxicityPrompts, or asks about evaluating this task. Reports Confidence Interval.

qhjqhj00 3eb07c5 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/toxicity-detection-eval commit 3eb07c57a5

Frequently asked questions

npx skillmds add qhjqhj00/toxicity-detection-eval