Pku Saferealf Eval

Probes an LLM's ability to generate safe and helpful responses by classifying harmful content across 19 distinct categories and 3 severity levels, while aligning with human preference rankings on Q-A-B triplets. Use when the user wants to benchmark on PKU-SafeRLHF, or asks about evaluating this task. Reports harm_category.

qhjqhj00 18dac9e 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/pku-saferealf-eval commit 18dac9e401

Frequently asked questions

npx skillmds add qhjqhj00/pku-saferealf-eval