Openrlhf Training

Train large language models (7B-70B+) with RLHF using PPO, GRPO, DPO, and other algorithms, accelerated by Ray and vLLM for distributed multi-GPU setups.

Orchestra Research Updated 10.4k repo stars

File contents

Orchestra-Research/AI-Research-SKILLs/tree/main/06-post-training/openrlhf commit 773a52944b

Frequently asked questions

npx skillmds@latest add orchestra-research/openrlhf-training