Group Fairness Reward Eval

Evaluates whether reward models assign equal average scores to high-quality responses across different demographic/occupational groups. It probes for systematic bias in how models rank expert-written abstracts based on the author's discipline. Use when the user wants to benchmark on arXiv Metadata (Curated), or asks about evaluating this task. Reports Normalized Maximum Group Difference.

qhjqhj00 350750f 2.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/group-fairness-reward-eval commit 350750f2aa

Frequently asked questions

npx skillmds add qhjqhj00/group-fairness-reward-eval