slime-rl-training

orchestra-research/slime-rl-training · Agent Skill (multi-file)

by Orchestra Research · bundle

Published · Last updated


Post-train LLMs with reinforcement learning using the slime framework, which integrates Megatron-LM for training and SGLang for rollout generation.

SKILL.md

Files

This skill is a package of 3 files. Install with the command above, or download the folder.

  • 📄SKILL.md entry
  • 📁references
  • 📄api-reference.md 11.6 KB
  • 📄troubleshooting.md 7.1 KB

Related

  1. slime-rl-training · lord1egypt
    Guides LLM post-training with RL using slime, a Megatron+SGLang framework for training GLM, Qwen, DeepSeek, and Llama models with GRPO, async, and multi-turn workflows.
    2
    repo stars
  2. nemo-mbridge-recipe-recommender · nvidia bundle
    Indexes Megatron Bridge recipes and recommends the best starting config based on model, GPU count, and training goal.
    2.2k
    repo stars
  3. miles-rl-training · orchestra-research bundle
    Train large-scale MoE models with FP8/INT4 low-precision RL, speculative decoding, and train-inference alignment using the miles framework.
    10.4k
    repo stars
  4. trl-training · huggingface
    Train and fine-tune transformer language models using TRL (Transformers Reinforcement Learning) with support for SFT, DPO, GRPO, KTO, RLOO, and reward model training via CLI commands.
    10.8k
    repo stars
  5. grpo-rl-training · orchestra-research bundle
    Expert guidance for implementing GRPO/RL fine-tuning with TRL for reasoning and task-specific model training.
    10.4k
    repo stars
  6. verl-rl-training · orchestra-research bundle
    Train LLMs with reinforcement learning using verl (Volcano Engine RL), supporting RLHF, GRPO, PPO, and other algorithms for scalable post-training with flexible infrastructure backends.
    10.4k
    repo stars

Frequently asked questions

How do I install the slime-rl-training skill?

Run npx skillmds add orchestra-research/slime-rl-training in your terminal (requires Node.js), paste this page's agent-chat prompt into Claude, Cursor, or any MCP-connected agent, or download the SKILL.md file and copy it into your agent's skills directory.

What does the slime-rl-training skill do?

Post-train LLMs with reinforcement learning using the slime framework, which integrates Megatron-LM for training and SGLang for rollout generation. It is listed under AI & ML, Model Training & Fine-tuning on SkillMD.

Is slime-rl-training safe to use?

SkillMD's automated safety review verdict for this skill is CAUTION. Independent scanners report: SkillSpector: PASS, Skill Scanner: PASS. Capability flags: makes network calls. SkillMD never runs a skill's scripts for you; review the SKILL.md before installing.

Which AI agents work with slime-rl-training?

This skill is tagged as working with Claude Code, Claude.ai, OpenAI Codex. SKILL.md is an open format, so most agents that read a skills directory can load it too.

Is slime-rl-training free to use?

Yes. Installing skills from SkillMD is free. This skill is licensed under MIT.

Who published slime-rl-training?

Orchestra Research (@orchestra-research) published this skill. Their other Agent Skills are listed on their SkillMD profile.