hallucinations-leaderboard-eval
The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models — Hong et al. (2024) (arXiv:2404.05904, 2024)
What this evaluates
This evaluation framework quantifies two distinct types of hallucinations in large language models: factuality (truthfulness of generated information) and faithfulness (consistency with input context or instructions). It probes model performance across 15 diverse knowledge-intensive tasks, including closed-book QA, summarization, reading comprehension, and fact-checking, using zero-shot and few-shot in-context prompts.
Datasets
- FEVER — total ?; splits: test (-1)
- FaithDial — total ?; splits: test (-1)
- NQ-open — total ?; splits: test (-1)
- TriviaQA — total ?; splits: test (-1)
- MemoTrap — total ?; splits: test (-1)
- TruthfulQA-MC2 — total ?; splits: test (-1)
- XSum — total ?; splits: test (-1)
Metrics
faithfulness score(primary) — range: percent- Measures the model's ability to retain fidelity to the given input context or specific instructions. Exact computation defined in Section 2.3.
factuality score(primary) — range: percent- Measures the accuracy and truthfulness of the generated information against ground-truth facts. Exact computation defined in Section 2.3.
Input / output format
Input: Task-specific prompts provided in zero-shot or few-shot in-context format, including questions, source documents, or instructions.
Output: Model-generated text or selected answer option.
Scoring recipe
for each task in evaluation_tasks:
predictions = model.generate(prompt)
factuality_scores[task] = compute_factuality(predictions, gold)
faithfulness_scores[task] = compute_faithfulness(predictions, context)
aggregate_factuality = mean(factuality_scores)
aggregate_faithfulness = mean(faithfulness_scores)
return aggregate_factuality, aggregate_faithfulness
Common pitfalls
- Scores are min-max normalized per task for heatmap visualization, but raw scores are reported in tables.
- Instruction fine-tuning improves faithfulness (instruction adherence) but does not consistently improve factuality, revealing a trade-off.
- Evaluation is conducted without training or fine-tuning the evaluated models.
Evidence (verbatim from paper)
we conducted a comprehensive analysis of the models and tasks introduced in Section [2]. In [Figure 2], we display the results of models for each task in the form of a heatmap. The value of each cell in the heatmap follows the metric of the corresponding task, and the dendrogram-shaped clusters are formed after applying min-max normalisation by task (y-axis) and model (x-axis). Table 1 shows a comparison of pre-trained models with their corresponding instruction fine-tuned variants across two metrics: faithfulness score and factuality score (Section [2.3]).
Citation
@misc{hong2024hallucinationsleaderboard,
title={The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models},
author={Hong et al. (2024)},
year={2024},
note={arXiv:2404.05904}
}
- arXiv: 2404.05904