AI Eval CI

Run AI agent and LLM evaluations in CI/CD pipelines — automated quality gates that fail the build when AI output quality drops. Use when someone asks to "test my AI agent", "add evals to CI", "catch prompt regressions", "compare models", "evaluate LLM output quality", "set up AI quality gates", or "benchmark my agent before deploying". Covers eval frameworks (Cobalt, Promptfoo, Braintrust), LLM-as-judge scoring, threshold-based assertions, and GitHub Actions integration.

eliferjunior Updated 0 repo stars

File contents

eliferjunior/Claude/tree/main/.claude/skills/ts-ai-eval-ci commit 0b96126770

Frequently asked questions

npx skillmds@latest add eliferjunior/ai-eval-ci