Safety Score

Quantifies implicit representational harms in pre-trained language models by measuring the disparity in language modeling probabilities between harmful and benign sentences targeting 13 marginalized demographics. It probes whether a model's internal likelihood estimates reflect toxic or stereotypical biases toward specific groups. Use when the user has predictions and gold and needs to compute safety score.

qhjqhj00 e65d4c1 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/safety-score commit e65d4c1ff6

Frequently asked questions

npx skillmds add qhjqhj00/safety-score