Falserject Eval

Evaluates an LLM's tendency to over-refuse benign prompts that merely appear harmful. It probes the model's ability to distinguish safe from unsafe contexts in controversial queries and provide helpful, context-aware responses instead of unnecessary refusals. Use when the user wants to benchmark on FalseReject, or asks about evaluating this task. Reports over-refusal.

qhjqhj00 481d94e 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/falserject-eval commit 481d94e113

Frequently asked questions

npx skillmds add qhjqhj00/falserject-eval