Peer Review Toxic Detection Eval

This benchmark evaluates the ability of models to detect toxic sentences within academic peer reviews, where toxicity manifests as subtle emotive, rhetorical, or unconstructive language rather than overt abuse. It measures alignment with human judgments and assesses whether models can revise toxic sentences while preserving the original critique. Use when the user wants to benchmark on Peer Review Toxic Detection Dataset, or asks about evaluating this task. Reports Cohen's Kappa.

qhjqhj00 e8ce09e 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/peer-review-toxic-detection-eval commit e8ce09e8a5

Frequently asked questions

npx skillmds add qhjqhj00/peer-review-toxic-detection-eval