Evaluation Reality Check

Design evaluation strategies that match task type, deployment setting, time, groups, and decision risk. Use when an agent needs a judgment-heavy data science workflow for design trustworthy model evaluation, including evidence review, local artifact inspection, risk classification, stakeholder-ready decisions, reproducibility, governance, or agent-to-agent handoff. Trigger for Codex, Claude, Gemini, Copilot, Cursor, Windsurf, Gravity, LangGraph, CrewAI, AutoGen, or local agents when this exact workflow is needed.

emily2040 Updated

File contents

emily2040/data-science-agent-skills/tree/main/data-science-agent-skills/evaluation-reality-check commit 5503121eab

Frequently asked questions

npx skillmds@latest add emily2040/evaluation-reality-check