# Multimodal Eval Suite Eval

> multimodal-eval-suite-eval

- Skill: `qhjqhj00/multimodal-eval-suite-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/multimodal-eval-suite-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/multimodal-eval-suite-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/multimodal-eval-suite-eval

---


# multimodal-eval-suite-eval

> AndesVL Technical Report: An Efficient Mobile-side Multimodal Large Language Model — Jin et al. (2025) (arXiv:2510.11496, 2025)

## What this evaluates

Evaluates multimodal large language models across six domains: reasoning & math, text-rich image understanding, multi-image comprehension, general VQA, hallucination mitigation, and multilingual capability. It probes the model's ability to process visual contexts, perform complex reasoning, extract text, and answer questions accurately across diverse real-world scenarios.

## Datasets

- **Multimodal Benchmark Suite (32 datasets)** — total ?; splits: test (-1)

## Metrics

- `Average accuracy / ANLS across 32 benchmarks` **(primary)** — range: percent
  - Arithmetic mean of per-benchmark scores (accuracy, ANLS, relaxed accuracy, or worst-case accuracy) across six domains and an overall aggregate.

## Input / output format

**Input**: Image(s) and text prompt/question per instance.

**Output**: Text answer generated by the model.

## Scoring recipe

```python
scores = []
for benchmark in benchmarks:
    preds = model.generate(image, prompt)
    gold = benchmark.answers
    score = compute_metric(preds, gold, metric=benchmark.metric) # accuracy, ANLS, etc.
    scores.append(score)
overall = sum(scores) / len(scores)
```

## Common pitfalls

- Scores for many baseline models are taken from original papers or the OpenCompass leaderboard rather than re-evaluated.
- Different benchmarks use different metrics (accuracy, ANLS, relaxed accuracy, worst-case accuracy) which are averaged directly without normalization.
- Evaluation is primarily conducted using VLMEvalKit, which may introduce framework-specific inference settings.

## Evidence (verbatim from paper)

> The accuracy results achieved from the model’s direct answer on its validation set are recorded. We compute the average scores, drawn from the models’ original papers or the OpenCompass leaderboard, to represent their capabilities across specific domains and overall.

## Citation

```bibtex
@misc{jin2025andesvl,
  title={AndesVL Technical Report: An Efficient Mobile-side Multimodal Large Language Model},
  author={Jin et al. (2025)},
  year={2025},
  note={arXiv:2510.11496}
}
```

- arXiv: 2510.11496

