# Cleva Eval

> cleva-eval

- Skill: `qhjqhj00/cleva-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/cleva-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/cleva-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/cleva-eval

---


# cleva-eval

> CLEVA: Chinese Language Models EVAluation Platform — Li et al. (2023) (arXiv:2308.04813, 2023)

## What this evaluates

Evaluates Chinese large language models across six key dimensions: accuracy, robustness, fairness, calibration, bias, and diversity. The platform uses standardized prompts and dynamic test set sampling to mitigate train-test contamination while comparing open-source and limited-access models.

## Datasets

- **CLEVA benchmark suite** — total ?; splits: test (-1); repo https://github.com/LaVi-Lab/CLEVA

## Metrics

- `Accuracy` **(primary)** — range: [0, 1]
  - Standard accuracy measuring the proportion of correct predictions or responses across tasks.
- `Robustness` — range: [0, 1]
  - Measures model performance stability under perturbations, following a similar trend to accuracy.
- `Fairness` — range: [0, 1]
  - Assesses performance on simplified-to-traditional Chinese conversion tasks.
- `Calibration (ECE-10)` — range: [0, 1]
  - Expected Calibration Error computed with 10 bins, following HELM conventions.
- `Bias` — range: [0, 1]
  - Measures gender bias in model outputs, noting that shorter outputs naturally correlate with lower measured bias.
- `Diversity` — range: [0, 1]
  - Inter-distinct metric comparing the diversity and innovation of generated expressions across models.

## Input / output format

**Input**: Standardized evaluation prompts/tasks from the CLEVA platform, with 6.43% of the test data sampled per leaderboard round.

**Output**: Model-generated text responses to the evaluation prompts.

## Scoring recipe

```python
def compute_leaderboard(predictions, gold, models):
    metrics = ['Accuracy', 'Robustness', 'Fairness', 'Calibration', 'Bias', 'Diversity']
    scores = {}
    for model in models:
        model_preds = predictions[model]
        model_scores = {}
        for m in metrics:
            model_scores[m] = compute_metric(model_preds, gold, m)
        scores[model] = mean(model_scores.values())
    return sort_by_mean_win_rate(scores, descending=True)
```

## Common pitfalls

- Train-test contamination is actively mitigated via dynamic sampling per round, so historical leaderboard scores may not reflect current model performance.
- Fairness evaluation relies on simplified-to-traditional Chinese conversion, which may unfairly penalize models lacking traditional Chinese pretraining data.
- Bias scoring is confounded by output length, as open-source models often produce shorter responses that naturally exhibit lower measured bias.

## Evidence (verbatim from paper)

> Results & Analysis. Figure 2 ranks all models by their mean win rates under different metrics. - Accuracy. It can be seen that GPT-4 (OpenAI, 2023) has the highest winning rate... - Calibration. We report ECE-10 (Kumar et al., 2019) following HELM. ... - Diversity. We choose inter-distinct to compare different models.

## Citation

```bibtex
@misc{li2023cleva,
  title={CLEVA: Chinese Language Models EVAluation Platform},
  author={Li et al. (2023)},
  year={2023},
  note={arXiv:2308.04813}
}
```

- arXiv: 2308.04813

