openrlhf-training

orchestra-research/openrlhf-training · Agent Skill (multi-file)

by Orchestra Research · bundle

Published · Last updated


Train large language models (7B-70B+) with RLHF using PPO, GRPO, DPO, and other algorithms, accelerated by Ray and vLLM for distributed multi-GPU setups.

SKILL.md

Files

This skill is a package of 5 files. Install with the command above, or download the folder.

  • 📄SKILL.md entry
  • 📁references
  • 📄algorithm-comparison.md 9.6 KB
  • 📄custom-rewards.md 15.5 KB
  • 📄hybrid-engine.md 7.1 KB
  • 📄multi-node-training.md 10.8 KB

Related

  1. huggingface-llm-trainer · huggingface bundle
    Train or fine-tune language and vision models using TRL or Unsloth on Hugging Face Jobs cloud infrastructure, with support for SFT, DPO, GRPO, and reward modeling, plus GGUF conversion for local deployment.
    10.8k
    repo stars
  2. verl-rl-training · orchestra-research bundle
    Train LLMs with reinforcement learning using verl (Volcano Engine RL), supporting RLHF, GRPO, PPO, and other algorithms for scalable post-training with flexible infrastructure backends.
    10.4k
    repo stars
  3. fine-tuning-with-trl · orchestra-research bundle
    Fine-tune and align language models using reinforcement learning with TRL, including SFT, DPO, PPO, GRPO, and reward model training.
    10.4k
    repo stars
  4. ray · majiayu000 bundle
    Scales Python ML workloads across clusters using Ray's distributed tasks, actors, data, and serving capabilities.
    567
    repo stars
  5. axolotl · orchestra-research bundle
    Provides expert guidance for fine-tuning LLMs with Axolotl, covering YAML configs, LoRA/QLoRA, DPO/KTO/ORPO/GRPO, and multimodal support.
    10.4k
    repo stars
  6. torchforge-rl-training · orchestra-research bundle
    Train reinforcement learning models using torchforge, Meta's PyTorch-native RL library for scalable, algorithm-focused experimentation with GRPO, DAPO, and custom loss functions.
    10.4k
    repo stars

Frequently asked questions

How do I install the openrlhf-training skill?

Run npx skillmds add orchestra-research/openrlhf-training in your terminal (requires Node.js), paste this page's agent-chat prompt into Claude, Cursor, or any MCP-connected agent, or download the SKILL.md file and copy it into your agent's skills directory.

What does the openrlhf-training skill do?

Train large language models (7B-70B+) with RLHF using PPO, GRPO, DPO, and other algorithms, accelerated by Ray and vLLM for distributed multi-GPU setups. It is listed under AI & ML, Coding & Dev Tools, DevOps & Infra, Containers & Kubernetes, Model Training & Fine-tuning on SkillMD.

Is openrlhf-training safe to use?

SkillMD's automated safety review verdict for this skill is CAUTION. Independent scanners report: SkillSpector: PASS, Skill Scanner: PASS. Capability flags: makes network calls, reads secrets. SkillMD never runs a skill's scripts for you; review the SKILL.md before installing.

Which AI agents work with openrlhf-training?

This skill is tagged as working with Claude Code, Claude.ai, OpenAI Codex. SKILL.md is an open format, so most agents that read a skills directory can load it too.

Is openrlhf-training free to use?

Yes. Installing skills from SkillMD is free. This skill is licensed under MIT.

Who published openrlhf-training?

Orchestra Research (@orchestra-research) published this skill. Their other Agent Skills are listed on their SkillMD profile.