# Videoscore2 Eval

> videoscore2-eval

- Skill: `qhjqhj00/videoscore2-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/videoscore2-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/videoscore2-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/videoscore2-eval

---


# videoscore2-eval

> VideoScore2: Think before You Score in Generative Video Evaluation — Xuan He et al. (arXiv:2509.22799, 2025)

## What this evaluates

This evaluation probes a model's ability to assess generative videos across three key dimensions: visual quality, text-to-video alignment, and physical/common-sense consistency. It measures how well automated scoring models align with human judgments on both in-domain and out-of-domain video benchmarks.

## Datasets

- **VideoGenReward Bench** — total 4691; splits: test (4691)
- **T2VQA-DB** — total 2000; splits: test (2000)
- **MJ-Bench-Video** — total 2170; splits: test (2170)
- **VideoPhy2-test** — total 3396; splits: test (3396)

## Metrics

- `accuracy` **(primary)** — range: percent
  - Percentage of correctly predicted video preferences or dimension scores compared to ground-truth human annotations. Calculated per benchmark after aligning scores to a unified 1–5 integer scale.

## Input / output format

**Input**: A generated video clip (and optionally the corresponding text prompt for alignment evaluation).

**Output**: Integer scores in the range [1, 5] for three dimensions: visual quality (v), text alignment (t), and physical consistency (p). For pairwise benchmarks, a preference label indicating which video is better.

## Scoring recipe

```python
def compute_accuracy(predictions, gold, benchmark):
    # 1. Align dimensions and rescale scores to unified [1,5] integer range
    pred_aligned = align_dimensions_and_rescale(predictions, benchmark)
    gold_aligned = align_dimensions_and_rescale(gold, benchmark)
    # 2. Compute preference or exact match
    if benchmark.type == 'pairwise':
        pred_score = pred_aligned.mean(axis=1)
        gold_score = gold_aligned.mean(axis=1)
        correct = (pred_score > gold_score) == (gold_score > pred_score)
    else:
        correct = (pred_aligned == gold_aligned).all(axis=1)
    return (correct.sum() / len(correct)) * 100
```

## Common pitfalls

- Different benchmarks use varying scoring scales ({0,1,2}, [1,100], [-2,2]) requiring careful rescaling before comparison.
- Some baselines output fewer dimensions than the three required, necessitating either broadcasting a single score or skipping missing dimensions.
- Pairwise preference benchmarks require averaging dimension scores to derive an overall preference, which can mask per-dimension errors.

## Evidence (verbatim from paper)

> VideoScore2 achieves superior accuracy (44.35 +5.94 on in-domain benchmarks) and strong out-of-domain generalization (50.37 +4.32 average), while providing transparent reasoning that enables effective reward modeling for Best-of-N video generation.

## Citation

```bibtex
@misc{he2025videoscore2,
  title={VideoScore2: Think before You Score in Generative Video Evaluation},
  author={Xuan He et al.},
  year={2025},
  note={arXiv:2509.22799}
}
```

- arXiv: 2509.22799

