Toxicity Analysis Eval

Measures the toxicity of text generated by language models conditioned on specific prompts, evaluating how alignment techniques like prompting and context distillation affect harmful content generation. Use when the user wants to benchmark on RealToxicityPrompts, or asks about evaluating this task. Reports mean toxicity score.

qhjqhj00 ddb4b46 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/toxicity-analysis-eval commit ddb4b46125

Frequently asked questions

npx skillmds add qhjqhj00/toxicity-analysis-eval