# Qgeval Eval

> qgeval-eval

- Skill: `qhjqhj00/qgeval-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/qgeval-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/qgeval-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/qgeval-eval

---


# qgeval-eval

> QGEval: Benchmarking Multi-dimensional Evaluation for Question Generation — Fu et al. (2024) (arXiv:2406.05707, 2024)

## What this evaluates

Evaluates the quality of generated questions across seven dimensions: fluency, clarity, conciseness, relevance, consistency, answerability, and answer consistency. It measures how well automatic metrics and LLMs align with human judgments on these dimensions.

## Datasets

- **QGEval** — total 3000; splits: test (3000); repo https://github.com/WeipingFu/QGEval

## Metrics

- `Pearson correlation` **(primary)** — range: [-1, 1]
  - Calculates the Pearson product-moment correlation coefficient between the vector of human annotation scores and the vector of automatic/LLM scores for each of the seven evaluation dimensions. Higher absolute values indicate stronger linear alignment with human judgments.

## Input / output format

**Input**: Passage, ground-truth answer, and generated question. For reference-based metrics, a reference question is also provided. Human annotation scores (1–5 Likert scale) per dimension, averaged over three annotators, serve as the gold standard.

**Output**: Per-dimension scores (1–5 scale) for each of the seven dimensions. Automatic metrics output a single composite score or per-dimension scores depending on the metric type.

## Scoring recipe

```python
for dim in ['fluency', 'clarity', 'conciseness', 'relevance', 'consistency', 'answerability', 'answer consistency']:
    human_scores = [get_human_score(q, dim) for q in questions]
    auto_scores = [get_auto_score(q, dim) for q in questions]
    pearson_r = pearsonr(human_scores, auto_scores)
    print(f'{dim}: {pearson_r:.3f}')
```

## Common pitfalls

- Ceiling effects in human scores for fluency, clarity, relevance, and consistency cause automatic metrics to show poor alignment despite high human ratings.
- Reference-based vs reference-free scoring types drastically change scores for metrics like BARTScore and GPTScore; the paper explicitly distinguishes ref-hypo and src-hypo configurations.
- The benchmark shows limited discriminative power among top-performing models on most dimensions, making it difficult to rank state-of-the-art QG models using standard t-tests.

## Evidence (verbatim from paper)

> We evaluate the agreement between the automatic metrics and human annotation scores by calculating the Pearson correlation over each dimension, results are shown in Table[5], with the three highest and lowest absolute coefficients bolded and underlined respectively.

## Citation

```bibtex
@misc{fu2024qgeval,
  title={QGEval: Benchmarking Multi-dimensional Evaluation for Question Generation},
  author={Fu et al. (2024)},
  year={2024},
  note={arXiv:2406.05707}
}
```

- arXiv: 2406.05707

