# Vdialogue Eval

> vdialogue-eval

- Skill: `qhjqhj00/vdialogue-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/vdialogue-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/vdialogue-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/vdialogue-eval

---


# vdialogue-eval

> VDialogUE: A Unified Evaluation Benchmark for Visually-grounded Dialogue — Yunshui Li et al. (2023) (arXiv:2309.07387, 2023)

## What this evaluates

Evaluates visually-grounded dialogue systems across five core tasks: multi-modal intent prediction, dialog retrieval (text-to-image and image-to-text), dialog state tracking, and response generation. It provides a unified hierarchical scoring metric to compare cross-task performance and generalization.

## Datasets

- **VisDial** — total ?; splits: test (-1)
- **PhotoChat** — total ?; splits: test (-1)
- **MMDialog** — total ?; splits: test (-1)
- **Image-Chat** — total ?; splits: test (-1)

## Metrics

- `VDscore` **(primary)** — range: [0, 1]
  - A hierarchical evaluation metric based on the Analytic Hierarchy Process (AHP) that aggregates task-specific performance scores into a single comprehensive score using weighted criteria.
- `R@1, R@5, R@10` — range: percent
  - Recall at K, measuring the proportion of instances where the ground-truth item appears within the top K ranked predictions.

## Input / output format

**Input**: Multi-modal dialogue context comprising text history and associated images, with task-specific prompts (e.g., candidate sets for retrieval, intent/state labels for prediction, or generation targets).

**Output**: Task-dependent: ranked list of images or text, predicted intent/state labels, or generated dialogue responses.

## Scoring recipe

```python
def compute_vdscore(task_scores):
    # AHP-based weighted aggregation of task-specific metrics
    return ahp_aggregate(task_scores)

def compute_recall_at_k(predictions, gold, k):
    return 1.0 if gold in predictions[:k] else 0.0
```

## Common pitfalls

- VisDial exhibits a distribution bias towards image content, causing models to ignore dialogue context.
- Annotator bias can create spurious causal links between dialogue context and output responses.
- Models struggle to differentiate correct images from visually similar candidates in retrieval tasks.
- Concatenating long dialogue history with short candidate answers equally degrades text retrieval performance.

## Evidence (verbatim from paper)

> Specifically, we found that our model also achieved consistent improvement in the comprehensive evaluation of VDscore.

## Citation

```bibtex
@misc{li2023vdialogue,
  title={VDialogUE: A Unified Evaluation Benchmark for Visually-grounded Dialogue},
  author={Yunshui Li et al. (2023)},
  year={2023},
  note={arXiv:2309.07387}
}
```

- arXiv: 2309.07387

