Polyglot Toxicity Prompts Eval

Evaluates the toxicity of LLM-generated continuations across 17 languages using naturally occurring prompts scraped from the web. It probes how model size, language resource availability, and instruction/preference tuning affect the generation of harmful content. Use when the user wants to benchmark on PolygloToxicityPrompts (PTP), or asks about evaluating this task. Reports AT.

qhjqhj00 e142897 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/polyglot-toxicity-prompts-eval commit e142897dde

Frequently asked questions

npx skillmds add qhjqhj00/polyglot-toxicity-prompts-eval