pope-nocaps-eval
A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models — Liqiang Jing et al. (2025) (arXiv:2505.01958, 2025)
What this evaluates
Tests object perception and hallucination on images without captions, evaluating whether LVLMs can ground object detection purely from visual input without textual priors.
Datasets
- POPE-NoCaps — total ?; splits: test (-1)
Metrics
Acc(primary) — range: [0, 1]- Accuracy: proportion of correct predictions out of total instances.
F1— range: [0, 1]- F1: harmonic mean of precision and recall for the positive class.
Input / output format
Input: Image paired with a yes/no question about object presence.
Output: Yes/No prediction.
Scoring recipe
def compute_metrics(preds, golds):
acc = sum(p == g for p, g in zip(preds, golds)) / len(golds)
tp = sum(1 for p, g in zip(preds, golds) if p == g == 'yes')
fp = sum(1 for p, g in zip(preds, golds) if p == 'yes' and g != 'yes')
fn = sum(1 for p, g in zip(preds, golds) if p != 'yes' and g == 'yes')
prec = tp / (tp + fp) if (tp + fp) > 0 else 0.0
rec = tp / (tp + fn) if (tp + fn) > 0 else 0.0
f1 = 2 * prec * rec / (prec + rec) if (prec + rec) > 0 else 0.0
return acc, f1
Common pitfalls
- Absence of captions removes textual grounding, making models more prone to language priors; evaluation must control for caption leakage.
- Binary yes/no scoring ignores confidence calibration, which is critical for hallucination detection.
Evidence (verbatim from paper)
Table 7: Performance of different methods on QA-FB15K.
| Method | Entity | | Relation | | | --- | | | | | | | Acc | F1 | Acc | F1 | | LLaVA-7B | 78.39 | 73.14 | 56.79 | 48.79 | ... Contrastive alignment objective is beneficial for cognition-based knowledge, as evidenced by the performance boost on QA-FB15K.
Citation
@misc{jing2025visualobjecthallucination,
title={A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models},
author={Liqiang Jing et al. (2025)},
year={2025},
note={arXiv:2505.01958}
}
- arXiv: 2505.01958