AI Model Evaluation

Systematically evaluate the performance, safety, and reliability of AI models and LLM applications. Master metric selection, LLM-as-a-judge, RAG evaluation, and human-in-the-loop processes.

hallucinaut 58cdfd9 4.4 KB Updated

File contents

hallucinaut/skills/tree/main/ai-model-evaluation commit 58cdfd9d53

Frequently asked questions

npx skillmds@latest add hallucinaut/ai-model-evaluation