multimodal-eval-suite-eval
AndesVL Technical Report: An Efficient Mobile-side Multimodal Large Language Model — Jin et al. (2025) (arXiv:2510.11496, 2025)
What this evaluates
Evaluates multimodal large language models across six domains: reasoning & math, text-rich image understanding, multi-image comprehension, general VQA, hallucination mitigation, and multilingual capability. It probes the model's ability to process visual contexts, perform complex reasoning, extract text, and answer questions accurately across diverse real-world scenarios.
Datasets
- Multimodal Benchmark Suite (32 datasets) — total ?; splits: test (-1)
Metrics
Average accuracy / ANLS across 32 benchmarks(primary) — range: percent- Arithmetic mean of per-benchmark scores (accuracy, ANLS, relaxed accuracy, or worst-case accuracy) across six domains and an overall aggregate.
Input / output format
Input: Image(s) and text prompt/question per instance.
Output: Text answer generated by the model.
Scoring recipe
scores = []
for benchmark in benchmarks:
preds = model.generate(image, prompt)
gold = benchmark.answers
score = compute_metric(preds, gold, metric=benchmark.metric) # accuracy, ANLS, etc.
scores.append(score)
overall = sum(scores) / len(scores)
Common pitfalls
- Scores for many baseline models are taken from original papers or the OpenCompass leaderboard rather than re-evaluated.
- Different benchmarks use different metrics (accuracy, ANLS, relaxed accuracy, worst-case accuracy) which are averaged directly without normalization.
- Evaluation is primarily conducted using VLMEvalKit, which may introduce framework-specific inference settings.
Evidence (verbatim from paper)
The accuracy results achieved from the model’s direct answer on its validation set are recorded. We compute the average scores, drawn from the models’ original papers or the OpenCompass leaderboard, to represent their capabilities across specific domains and overall.
Citation
@misc{jin2025andesvl,
title={AndesVL Technical Report: An Efficient Mobile-side Multimodal Large Language Model},
author={Jin et al. (2025)},
year={2025},
note={arXiv:2510.11496}
}
- arXiv: 2510.11496