LLM Evaluation

Implement comprehensive evaluation strategies for LLM applications using automated metrics, LLM-as-judge, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, comparing prompts/models, or establishing evaluation frameworks. Covers RAGAS for RAG pipelines, evals-as-code CI/CD integration, and modern 2025/2026 practices including structured output evaluation and agentic task success measurement.

ckorhonen d1b533c 21.7 KB Updated

File contents

ckorhonen/claude-skills/tree/main/skills/llm-evaluation commit d1b533c3c7

Frequently asked questions

npx skillmds@latest add ckorhonen/llm-evaluation