Run Evals

Prepare the environment and run the LLM-driven agent evals (e2e/agent-evals/) against a chosen sim-use binary. Use when the user runs `/run-evals` or asks to "run the agent evals", "run the LLM-driven tests", "eval the skill", or wants pre-release confidence that an agent reading the bundled skill still picks the right verbs. Costs real `claude -p` API calls — always confirm before spending.

lycorp-jp 9fd14de 5.3 KB Updated

File contents

lycorp-jp/sim-use/tree/main/.agents/skills/run-evals commit 9fd14de13a

Frequently asked questions

npx skillmds@latest add lycorp-jp/run-evals