Evaluating Llms

Evaluate LLM systems using automated metrics, LLM-as-judge, and benchmarks. Use when testing prompt quality, validating RAG pipelines, measuring safety (hallucinations, bias), or comparing models for production deployment. Use when this capability is needed.

tomevault-io 947b5da 2 files · 19.1 KB Updated

File contents

tomevault-io/skills-registry/tree/main/ancoleman--ai-design-components--evaluating-llms commit 947b5da529

Frequently asked questions

npx skillmds@latest add tomevault-io/evaluating-llms