mm-upd-eval
Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models — Miyai et al. (2024) (arXiv:2403.20331, 2024)
What this evaluates
Evaluates Large Multimodal Models' ability to detect unsolvable visual-question problems and appropriately refuse to answer. It specifically probes three failure modes: absent answers, incompatible answer sets, and image-question mismatches.
Datasets
- MM-UPD Bench — total ?; splits: test (-1); repo https://github.com/AtsuMiyai/UPD
Metrics
Dual accuracy(primary) — range: [0, 1]- The combined accuracy of correctly answering standard solvable questions and correctly identifying/refusing unsolvable questions. Calculated as the proportion of instances where the model either provides the correct answer (for solvable) or correctly selects a refusal option (for unsolvable).
Input / output format
Input: An image paired with a multiple-choice question. For unsolvable instances, the question or options are modified to create absent answers, incompatible answer sets, or image-question mismatches.
Output: A single selected option from the provided choices, or an explicit refusal (e.g., 'None of the above').
Scoring recipe
def compute_dual_accuracy(predictions, golds, is_unsolvable):
correct = 0
for pred, gold, unsolvable in zip(predictions, golds, is_unsolvable):
if unsolvable:
if pred == gold: # gold is the refusal option
correct += 1
else:
if pred == gold:
correct += 1
return correct / len(predictions)
Common pitfalls
- High accuracy on standard benchmarks like MMBench does not predict performance on unsolvable detection; correlation is very low.
- Refusal capability is highly sensitive to model size and training data; open-source models often lack refusal alignment compared to closed-source ones.
- Prompting techniques like Chain-of-Thought or self-reflection can improve refusal for some models but may degrade performance for others depending on their specific bottlenecks.
Evidence (verbatim from paper)
Table 1 presents the overall Dual accuracies. Also, we show the Standard and UPD accuracies for some LMMs in Fig. 3. In Fig. 4, we show the radar charts of InternVL2-40B and GPT-4o for ability-wise fine-grained analysis.
Citation
@misc{miyai2024unsolvable,
title={Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models},
author={Miyai et al. (2024)},
year={2024},
note={arXiv:2403.20331}
}
- arXiv: 2403.20331