Run LLM Evals

Runs checklists and workflows for designing, operating, and improving LLM evals — domain harnesses, agent/voice evals, LLM judges, metrics workshops, product flywheels, dynamic search eval, enterprise conversation intelligence. Use when the user says "evals", "LLM judge", "eval harness", "regression on prompts", "agent eval", "eval dataset", or needs to ship with measurement before prod.

hiteshbandhu Updated

File contents

hiteshbandhu/skills-i-use/tree/main/skills/ai-engineer-talks/run-llm-evals commit 9db5759794

Frequently asked questions

npx skillmds@latest add hiteshbandhu/run-llm-evals