Native Agent Evals

Evaluate or audit Codex native Multi-Agent V2 behavior with small reproducible probes and structural rollout evidence. Use for V2 surface checks, direct-tool lifecycle behavior, bounded fanout, interruption semantics, resume behavior, or skill-invocation evidence.

Kbediako f58a573 5 files · 333.4 KB Updated

File contents

Kbediako/evergreen-codex-skills/tree/main/skills/native-agent-evals commit f58a57305d

Frequently asked questions

npx skillmds@latest add kbediako/native-agent-evals