miles-rl-training

orchestra-research/miles-rl-training · Agent Skill (multi-file)

by Orchestra Research · bundle

Published · Last updated


Train large-scale MoE models with FP8/INT4 low-precision RL, speculative decoding, and train-inference alignment using the miles framework.

SKILL.md

Files

This skill is a package of 3 files. Install with the command above, or download the folder.

  • 📄SKILL.md entry
  • 📁references
  • 📄api-reference.md 4.0 KB
  • 📄troubleshooting.md 5.7 KB

Related

  1. slime-rl-training · orchestra-research bundle
    Post-train LLMs with reinforcement learning using the slime framework, which integrates Megatron-LM for training and SGLang for rollout generation.
    10.4k
    repo stars
  2. nemo-mbridge-recipe-recommender · nvidia bundle
    Indexes Megatron Bridge recipes and recommends the best starting config based on model, GPU count, and training goal.
    2.2k
    repo stars
  3. slime-rl-training · lord1egypt
    Guides LLM post-training with RL using slime, a Megatron+SGLang framework for training GLM, Qwen, DeepSeek, and Llama models with GRPO, async, and multi-turn workflows.
    2
    repo stars
  4. nemo-automodel-model-onboarding · nvidia bundle
    Guides implementation of new model architectures in NeMo AutoModel through five phases: discovery, implementation, registration, validation, and testing.
    2.2k
    repo stars
  5. stable-baselines3 · k-dense-ai bundle
    Train reinforcement learning agents using PPO, SAC, DQN, TD3, DDPG, and A2C algorithms with a scikit-learn-like API. Supports custom Gymnasium environments, vectorized environments, callbacks, and model persistence.
    30.2k
    repo stars
  6. trl-training · huggingface
    Train and fine-tune transformer language models using TRL (Transformers Reinforcement Learning) with support for SFT, DPO, GRPO, KTO, RLOO, and reward model training via CLI commands.
    10.8k
    repo stars

Frequently asked questions

How do I install the miles-rl-training skill?

Run npx skillmds add orchestra-research/miles-rl-training in your terminal (requires Node.js), paste this page's agent-chat prompt into Claude, Cursor, or any MCP-connected agent, or download the SKILL.md file and copy it into your agent's skills directory.

What does the miles-rl-training skill do?

Train large-scale MoE models with FP8/INT4 low-precision RL, speculative decoding, and train-inference alignment using the miles framework. It is listed under AI & ML, Model Training & Fine-tuning on SkillMD.

Is miles-rl-training safe to use?

SkillMD's automated safety review verdict for this skill is CAUTION. Independent scanners report: SkillSpector: PASS, Skill Scanner: PASS. Capability flags: makes network calls. SkillMD never runs a skill's scripts for you; review the SKILL.md before installing.

Which AI agents work with miles-rl-training?

This skill is tagged as working with Claude Code, Claude.ai, OpenAI Codex. SKILL.md is an open format, so most agents that read a skills directory can load it too.

Is miles-rl-training free to use?

Yes. Installing skills from SkillMD is free. This skill is licensed under MIT.

Who published miles-rl-training?

Orchestra Research (@orchestra-research) published this skill. Their other Agent Skills are listed on their SkillMD profile.