LLM Evaluation
Evaluate LLM response quality with automatic metrics and LLM-as-judge.
Quick Start
from rouge import Rouge
def evaluate_summary(reference, candidate):
rouge = Rouge()
scores = rouge.get_scores(candidate, reference)
return scores[0]['rouge-l']['f']
When to Use
- ✅ Testing LLM quality
- ✅ Comparing different models
- ❌ Not for evaluating classifier accuracy
Step-by-Step Instructions
- Prepare test dataset with reference answers
- Generate responses with model under test
- Calculate metrics (BLEU, ROUGE, BERTScore)
- Conduct LLM-as-judge evaluation
Dependencies
pip install rouge-score bert-score openai
Examples
Input: reference="Hello", candidate="Hello!" → Output: ROUGE-L F1 = 0.95
Resources
Validation
- Metrics calculated correctly
- High correlation with human judgment
- Reports generated automatically