# LLM Eval

> Evaluates LLM performance using BLEU, ROUGE metrics and LLM-as-judge. Use for model testing.

- Skill: `ssrjkk/llm-eval` (Agent Skill, multi-file: 2 files)
- Install (CLI): `npx skillmds@latest add ssrjkk/llm-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/ssrjkk/llm-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- Author: ssrjkk (https://skillmd.com/u/ssrjkk)
- Updated: 2026-08-19
- Page: https://skillmd.com/skills/ssrjkk/llm-eval

---

# LLM Evaluation

> Evaluate LLM response quality with automatic metrics and LLM-as-judge.

## Quick Start
```python
from rouge import Rouge

def evaluate_summary(reference, candidate):
    rouge = Rouge()
    scores = rouge.get_scores(candidate, reference)
    return scores[0]['rouge-l']['f']
```

## When to Use
- ✅ Testing LLM quality
- ✅ Comparing different models
- ❌ Not for evaluating classifier accuracy

## Step-by-Step Instructions
1. Prepare test dataset with reference answers
2. Generate responses with model under test
3. Calculate metrics (BLEU, ROUGE, BERTScore)
4. Conduct LLM-as-judge evaluation

## Dependencies
```bash
pip install rouge-score bert-score openai
```

## Examples
Input: reference="Hello", candidate="Hello!" → Output: ROUGE-L F1 = 0.95

## Resources
- [BLEU Score](https://en.wikipedia.org/wiki/BLEU)
- [Examples](./examples/)

## Validation
1. Metrics calculated correctly
2. High correlation with human judgment
3. Reports generated automatically

