Redrft Eval

Evaluates reinforcement fine-tuning methods for red-teaming LLMs by measuring the toxicity and diversity of generated adversarial prompts across toxic continuation and instruction-following tasks. Use when the user wants to benchmark on toxic continuation, instruction following, or asks about evaluating this task. Reports cumulative toxicity-diversity score.

qhjqhj00 1b048d3 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/redrft-eval commit 1b048d38b7

Frequently asked questions

npx skillmds add qhjqhj00/redrft-eval