Eval Engine

Use this skill when a user supplies an AI feature spec, PRD, PMOS/task contract, existing eval suite, traces, outputs, state evidence, or release question and wants to define good, create or run evals, grade outcomes, agent trajectories, end-to-end system checkpoints, or promised memory behavior, add deterministic or model graders, calibrate an LLM judge against human goldens, compare repeated trials, inspect failures, debug evaluation evidence, or make a CI release decision. Produce one provider-neutral AI Evals for PMs suite through the stable pm-verifier runtime, with binary gates, gradual rubrics, provenance, operational metrics, JSON results, and a Markdown report. Also use for migrating pm-evals or Evals-pass-1 work. Do not use for ordinary non-AI unit testing, generating the AI application itself, a one-off opinion on one output without a quality contract, or live production monitoring.

Abhillashjadhav Updated

File contents

Abhillashjadhav/AI-PM-essential-skills/tree/main/pm-verifier/skills/eval-engine commit a0db7f904e

Frequently asked questions

npx skillmds@latest add abhillashjadhav/eval-engine-2