Ceb Fairness Eval

This benchmark evaluates how large language models exhibit social bias across different tasks, bias types, and social groups. It systematically measures bias through direct classification tasks and indirect text generation tasks, using standardized metrics to enable cross-dataset fairness comparisons. Use when the user wants to benchmark on CEB, or asks about evaluating this task. Reports Micro-F1.

qhjqhj00 0d3e717 4.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/ceb-fairness-eval commit 0d3e717da2

Frequently asked questions

npx skillmds add qhjqhj00/ceb-fairness-eval