Nemotron Cascade Rl

Train language models through sequential, domain-wise RL stages (RLHF → Instruction-Following → Math → Code → SWE) without catastrophic forgetting. Exploit policy-dependent training data distribution where previous behaviors persist when reward-relevant. 14B model surpasses DeepSeek-R1-0528 (671B) on LiveCodeBench.

adu2021 d0091e6 3.3 KB Updated

File contents

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.2-claude-opus-4.6/nemotron-cascade-rl commit d0091e63ea

Frequently asked questions

npx skillmds@latest add adu2021/nemotron-cascade-rl