Bradley Terry

Evaluates the stability and reliability of global pointwise scores (accuracy, AUC, F1) versus pairwise Bradley-Terry rankings for ordering NLP models across classification and text generation tasks. Use when the user has predictions and gold and needs to compute Bradley-Terry.

qhjqhj00 ddd9872 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/bradley-terry commit ddd9872204

Frequently asked questions

npx skillmds add qhjqhj00/bradley-terry