# Do Bench Eval

> do-bench-eval

- Skill: `qhjqhj00/do-bench-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/do-bench-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/do-bench-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/do-bench-eval

---


# do-bench-eval

> DO-Bench: An Attributable Benchmark for Diagnosing Object Hallucination in Vision-Language Models — Wang et al. (2026) (arXiv:2604.22822, 2026)

## What this evaluates

DO-Bench probes object hallucination in vision-language models by disentangling it into two distinct failure mechanisms: prior-dominated (textual priors overriding visual evidence) and perception-limited (weak visual grounding causing false denials). It uses controlled within-image interventions to measure how models respond to strengthened contextual priors and concentrated visual evidence, revealing heterogeneous failure modes that aggregate accuracy masks.

## Datasets

- **DO-Bench** — total 1240; splits: test (1240)

## Metrics

- `Acc` — range: percent
  - Standard accuracy over the unified yes/no answering protocol. Calculated as the fraction of instances where the normalized model prediction matches the ground truth label.
- `F1` — range: percent
  - F1 score computed over the yes/no predictions, balancing precision and recall across the dataset.
- `PerceptionAbility` **(primary)** — range: percent
  - Measures recognition recoverability under enhanced visual evidence at the baseline prompt (A0). Aggregates false-negative rates under Cluster and Crop views; higher values indicate stronger perceptual grounding and recoverability when localized evidence is provided.
- `PriorRobust` **(primary)** — range: percent
  - Aggregates robustness against prior-driven denial (A-series) and prior-driven completion (B-series). Computed via AUC aggregation over false-negative and false-positive rates across progressively strengthened contextual prior levels (A0–A3 and B0–B3). Higher values indicate greater resistance to textual prior override.

## Input / output format

**Input**: An image paired with a text prompt. The prompt varies by experimental condition: prior strength levels (A0–A3 for present objects, B0–B3 for absent objects) and visual evidence concentration (Full, Cluster, or Crop views). All instances follow a unified yes/no answering protocol.

**Output**: A deterministic yes/no answer (normalized to match ground truth labels).

## Scoring recipe

```python
def score(predictions, golds, conditions):
    correct = sum(1 for p, g in zip(predictions, golds) if p == g)
    acc = correct / len(golds) * 100
    
    # PerceptionAbility: A-series at baseline prompt (A0)
    fn_cluster = sum(1 for p, g in zip(predictions, golds) if p != g and g == 'Yes' and conditions == 'Cluster')
    fn_crop = sum(1 for p, g in zip(predictions, golds) if p != g and g == 'Yes' and conditions == 'Crop')
    perception_ability = 100 - (fn_cluster + fn_crop) / (2 * total_A0) * 100
    
    # PriorRobust: AUC of error rates across prior levels A0-A3
    fn_rates = [count_false_negatives(level) for level in ['A0','A1','A2','A3']]
    fp_rates = [count_false_positives(level) for level in ['B0','B1','B2','B3']]
    prior_robust = 100 - (auc(fn_rates) + auc(fp_rates)) / 2 * 100
    return acc, perception_ability, prior_robust
```

## Common pitfalls

- Relying solely on aggregate accuracy or F1 conflates qualitatively different failure modes; a model can have high accuracy but low PriorRobust, indicating severe susceptibility to textual priors.
- Assuming model responses must be strictly monotonic across prior strength or evidence concentration levels; the benchmark explicitly allows non-monotonicity and uses AUC aggregation to remain robust to discrete decoding effects.
- Treating perceptual recoverability and prior robustness as interchangeable; they are orthogonal dimensions, with evidence concentration primarily shifting prior-response curves vertically rather than altering their shape.

## Evidence (verbatim from paper)

> PerceptionAbility measures recognition under enhanced visual evidence under the baseline prompt (A0). PriorRobust aggregates robustness against prior-driven denial (A-series) and prior-driven completion (B-series). All values are in %.

## Citation

```bibtex
@misc{wang2026dobench,
  title={DO-Bench: An Attributable Benchmark for Diagnosing Object Hallucination in Vision-Language Models},
  author={Wang et al. (2026)},
  year={2026},
  note={arXiv:2604.22822}
}
```

- arXiv: 2604.22822

