Beavertails Moderation Eval

This evaluation probes the safety moderation and context-comprehension capabilities of external text moderation APIs. It measures how well automated systems align with human and expert-preference labels when assessing harmfulness across specific risk categories in QA pairs. Use when the user wants to benchmark on BeaverTails Evaluation Dataset, or asks about evaluating this task. Reports agreement.

qhjqhj00 151ad5d 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/beavertails-moderation-eval commit 151ad5d95b

Frequently asked questions

npx skillmds add qhjqhj00/beavertails-moderation-eval