# Visres Bench Eval

> visres-bench-eval

- Skill: `qhjqhj00/visres-bench-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/visres-bench-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/visres-bench-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/visres-bench-eval

---


# visres-bench-eval

> VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs — Malagurski Törtei et al. (2025) (arXiv:2512.21194, 2025)

## What this evaluates

This benchmark evaluates the visual reasoning capabilities of vision-language models across a perceptual-to-reasoning continuum. It isolates three levels of difficulty: basic perceptual grounding under transformations, single-attribute reasoning (color, count, orientation), and multi-attribute compositional reasoning. The setup tests whether models rely on genuine visual abstraction or fall back to linguistic priors when faced with naturalistic perturbations and rule-based inference.

## Datasets

- **VisRes Bench** — total ?; splits: test (-1)

## Metrics

- `accuracy` **(primary)** — range: percent
  - Percentage of correctly answered samples out of the total number of evaluated instances. A response is counted as correct only if the model outputs a definitive answer choice (A–D) without looping or exceeding context limits.

## Input / output format

**Input**: A main image paired with four candidate options (A–D). Each instance includes metadata specifying the subtask type, distractor generation strategy, augmentation parameters, ground-truth coordinates, and the correct answer index. Models are prompted with either a generic prompt (minimal guidance) or a guided prompt (specifying the visual attribute and relational rule).

**Output**: A single definitive answer choice (A, B, C, or D). Responses that loop in reasoning, repeat thinking, or exceed the 32k context limit without concluding are treated as incorrect.

## Scoring recipe

```python
def compute_accuracy(predictions, gold_labels):
    correct = 0
    total = len(predictions)
    for pred, gold in zip(predictions, gold_labels):
        if pred is not None and pred == gold:
            correct += 1
    return (correct / total) * 100
```

## Common pitfalls

- Models may enter reasoning loops or exceed the 32k context window without outputting a final answer; these must be explicitly parsed and marked incorrect rather than ignored.
- Performance varies significantly between guided and generic prompts; results are not directly comparable across prompt variants without controlling for the reasoning mode (e.g., 'thinking mode' enabled).
- Perceptual perturbations (occlusion, blur, edge detection) drastically reduce accuracy, so evaluating on unperturbed images alone overestimates true visual reasoning capability.

## Evidence (verbatim from paper)

> Performance is measured using accuracy, and results are reported both per subtask and in aggregate to compare model performance at different levels of granularity.

## Citation

```bibtex
@misc{malagurskitortei2025visresbench,
  title={VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs},
  author={Malagurski Törtei et al. (2025)},
  year={2025},
  note={arXiv:2512.21194}
}
```

- arXiv: 2512.21194

