Eval

Measure non-deterministic behavior — LLM features, agents, prompts, or a skill itself — with repeatable evals instead of one-shot checks. Use when building or tuning AI/LLM functionality (ranking, extraction, generation, agent loops), when a feature could pass once by luck, or when validating that a prompt or skill actually changes behavior.

vasu-devs 4a47e13 4.3 KB Updated

File contents

vasu-devs/forge/tree/main/skills/eval commit 4a47e13f36

Frequently asked questions

npx skillmds@latest add vasu-devs/eval