Tovo Consensus Eval

This evaluation probes a model's ability to classify text content according to a user-defined toxicity taxonomy. It measures how closely the model's predictions align with gold labels generated through a multi-model voting process, and tests generalization to out-of-domain categories. Use when the user wants to benchmark on ToVo, or asks about evaluating this task. Reports consensus rate.

qhjqhj00 20be702 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/tovo-consensus-eval commit 20be702365

Frequently asked questions

npx skillmds add qhjqhj00/tovo-consensus-eval