Mars Rl Multi Agent Reasoning

Train multi-agent reasoning systems with decoupled reward signals and pipeline parallelism—enable specialized Solver/Verifier/Corrector agents to iteratively refine solutions without waiting for full trajectories, handling extended reasoning up to 320K tokens.

adu2021 33a2d54 9.9 KB Updated

File contents

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.2-claude-opus-4.6/mars-rl-multi-agent-reasoning commit 33a2d54f09

Frequently asked questions

npx skillmds@latest add adu2021/mars-rl-multi-agent-reasoning