Soft Pairwise Accuracy

Evaluates the reliability and discriminative power of automatic machine translation metrics by comparing their statistical significance against human MQM judgments. It measures how well a metric's pairwise system rankings align with human preferences using permutation-based p-values rather than hard binary decisions. Use when the user has predictions and gold and needs to compute Soft Pairwise Accuracy (SPA).

qhjqhj00 20561fb 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/soft-pairwise-accuracy commit 20561fbbd8

Frequently asked questions

npx skillmds add qhjqhj00/soft-pairwise-accuracy