Senior Eval Engineer

Use when designing an eval set or eval harness for an LLM app, agent, RAG pipeline, classifier, or generative output; building a gold set; configuring an LLM as judge with a rubric; calibrating a judge against human raters; designing slice metrics; wiring a regression suite into CI; running a vibe check with rigor; choosing between exact match, BLEU, ROUGE, BERTScore, faithfulness, groundedness, or retrieval recall at K; computing inter rater agreement (Cohen kappa, Krippendorff alpha); auditing judge drift; or reporting eval deltas vs a baseline. Triggers: eval, evaluation, LLM eval, judge, LLM as judge, gold set, holdout, regression suite, agent eval, rubric, calibration, inter rater agreement, Cohen kappa, vibe check. Produces eval task specs, gold set construction plans, judge configurations, harness run reports, regression gate policies. Not for the model itself (training, serving), see senior-ml-engineer; not for the prompt as a product, see senior-llm-app-engineer.

iamdemetris 77609a7 20.8 KB Updated

File contents

iamdemetris/lude-kit/tree/main/skills/personas/senior-eval-engineer commit 77609a7b78

Frequently asked questions

npx skillmds@latest add iamdemetris/senior-eval-engineer