Much Eval

Evaluates the ability of logit-based uncertainty quantification (UQ) methods to predict claim-level hallucination (factuality) in multilingual LLM outputs. It measures how well token-level confidence scores, when aggregated, correlate with ground-truth factuality labels across different languages and model configurations. Use when the user wants to benchmark on MUCH, or asks about evaluating this task. Reports ROC-AUC.

qhjqhj00 fc063c9 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/much-eval commit fc063c9efc

Frequently asked questions

npx skillmds add qhjqhj00/much-eval