Prepare Verifier Handoff

Use when a workload must learn multi-step behavior through hosted reinforcement-learning (RL) training that local optimization cannot deliver, and needs to become partner-ready. "My agent needs RL", "can we train this policy", "package my environment for a training partner", "is this workload ready for RL". Decides first, then authors the trainable environment, then packages it — never trains here.

understudylabs Updated

File contents

understudylabs/understudy-agent-tools/tree/main/skills/prepare-verifier-handoff commit dfd75181cc

Frequently asked questions

npx skillmds@latest add understudylabs/prepare-verifier-handoff