Average Per Token Log Prob

Evaluates language models on multiple-choice or candidate-selection downstream tasks by scoring candidate answers based on their likelihood under the model. It measures how well the model assigns high probability to the correct answer among a set of options. Use when the user has predictions and gold and needs to compute average_per_token_log_prob.

qhjqhj00 3f950a5 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/average_per_token_log_prob commit 3f950a548d

Frequently asked questions

npx skillmds add qhjqhj00/average-per-token-log-prob