verl-rl-training

orchestra-research/verl-rl-training · Agent Skill (multi-file)

by Orchestra Research · bundle

Published · Last updated


Train LLMs with reinforcement learning using verl (Volcano Engine RL), supporting RLHF, GRPO, PPO, and other algorithms for scalable post-training with flexible infrastructure backends.

SKILL.md

Files

This skill is a package of 3 files. Install with the command above, or download the folder.

  • 📄SKILL.md entry
  • 📁references
  • 📄api-reference.md 6.8 KB
  • 📄troubleshooting.md 6.6 KB

Related

  1. openrlhf-training · orchestra-research bundle
    Train large language models (7B-70B+) with RLHF using PPO, GRPO, DPO, and other algorithms, accelerated by Ray and vLLM for distributed multi-GPU setups.
    10.4k
    repo stars
  2. fine-tuning-with-trl · orchestra-research bundle
    Fine-tune and align language models using reinforcement learning with TRL, including SFT, DPO, PPO, GRPO, and reward model training.
    10.4k
    repo stars
  3. stable-baselines3 · k-dense-ai bundle
    Train reinforcement learning agents using PPO, SAC, DQN, TD3, DDPG, and A2C algorithms with a scikit-learn-like API. Supports custom Gymnasium environments, vectorized environments, callbacks, and model persistence.
    30.2k
    repo stars
  4. trl-training · huggingface
    Train and fine-tune transformer language models using TRL (Transformers Reinforcement Learning) with support for SFT, DPO, GRPO, KTO, RLOO, and reward model training via CLI commands.
    10.8k
    repo stars
  5. pufferlib · k-dense-ai bundle
    Train reinforcement learning agents at millions of steps per second using optimized PPO, vectorized environments, and multi-agent support.
    30.2k
    repo stars
  6. slime-rl-training · lord1egypt
    Guides LLM post-training with RL using slime, a Megatron+SGLang framework for training GLM, Qwen, DeepSeek, and Llama models with GRPO, async, and multi-turn workflows.
    2
    repo stars

Frequently asked questions

How do I install the verl-rl-training skill?

Run npx skillmds add orchestra-research/verl-rl-training in your terminal (requires Node.js), paste this page's agent-chat prompt into Claude, Cursor, or any MCP-connected agent, or download the SKILL.md file and copy it into your agent's skills directory.

What does the verl-rl-training skill do?

Train LLMs with reinforcement learning using verl (Volcano Engine RL), supporting RLHF, GRPO, PPO, and other algorithms for scalable post-training with flexible infrastructure backends. It is listed under AI & ML, Model Training & Fine-tuning on SkillMD.

Is verl-rl-training safe to use?

SkillMD's automated safety review verdict for this skill is CAUTION. Independent scanners report: SkillSpector: PASS, Skill Scanner: PASS. Capability flags: makes network calls. SkillMD never runs a skill's scripts for you; review the SKILL.md before installing.

Which AI agents work with verl-rl-training?

This skill is tagged as working with Claude Code, Claude.ai, OpenAI Codex. SKILL.md is an open format, so most agents that read a skills directory can load it too.

Is verl-rl-training free to use?

Yes. Installing skills from SkillMD is free. This skill is licensed under MIT.

Who published verl-rl-training?

Orchestra Research (@orchestra-research) published this skill. Their other Agent Skills are listed on their SkillMD profile.