Fairness And Downstream Eval

Evaluates demographic fairness and downstream NLU task performance of language models. It measures bias across gender, race, and age using established fairness benchmarks, and verifies that fairness interventions do not degrade accuracy on standard classification and regression tasks. Use when the user wants to benchmark on HolisticBias, WEAT/SEAT, CrowS-Pairs, GLUE, or asks about evaluating this task. Reports Fairscore.

qhjqhj00 e4316e2 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/fairness-and-downstream-eval commit e4316e2a00

Frequently asked questions

npx skillmds add qhjqhj00/fairness-and-downstream-eval