fine-tuning-with-trl

orchestra-research/fine-tuning-with-trl · Agent Skill (multi-file)

by Orchestra Research · bundle

Published · Last updated


Fine-tune and align language models using reinforcement learning with TRL, including SFT, DPO, PPO, GRPO, and reward model training.

SKILL.md

Files

This skill is a package of 5 files. Install with the command above, or download the folder.

  • 📄SKILL.md entry
  • 📁references
  • 📄dpo-variants.md 4.2 KB
  • 📄online-rl.md 1.9 KB
  • 📄reward-modeling.md 2.5 KB
  • 📄sft-training.md 3.2 KB

Related

  1. trl-training · huggingface
    Train and fine-tune transformer language models using TRL (Transformers Reinforcement Learning) with support for SFT, DPO, GRPO, KTO, RLOO, and reward model training via CLI commands.
    10.8k
    repo stars
  2. huggingface-llm-trainer · huggingface bundle
    Train or fine-tune language and vision models using TRL or Unsloth on Hugging Face Jobs cloud infrastructure, with support for SFT, DPO, GRPO, and reward modeling, plus GGUF conversion for local deployment.
    10.8k
    repo stars
  3. grpo-rl-training · orchestra-research bundle
    Expert guidance for implementing GRPO/RL fine-tuning with TRL for reasoning and task-specific model training.
    10.4k
    repo stars
  4. gemma-trainer · google-gemma bundle
    Fine-tune Gemma models locally using QLoRA, Unsloth, or TRL for SFT, DPO, and reward modeling, with dataset preparation and conversion to GGUF or LiteRT-LM.
    3
    installs
  5. openrlhf-training · orchestra-research bundle
    Train large language models (7B-70B+) with RLHF using PPO, GRPO, DPO, and other algorithms, accelerated by Ray and vLLM for distributed multi-GPU setups.
    10.4k
    repo stars
  6. verl-rl-training · orchestra-research bundle
    Train LLMs with reinforcement learning using verl (Volcano Engine RL), supporting RLHF, GRPO, PPO, and other algorithms for scalable post-training with flexible infrastructure backends.
    10.4k
    repo stars

Frequently asked questions

How do I install the fine-tuning-with-trl skill?

Run npx skillmds add orchestra-research/fine-tuning-with-trl in your terminal (requires Node.js), paste this page's agent-chat prompt into Claude, Cursor, or any MCP-connected agent, or download the SKILL.md file and copy it into your agent's skills directory.

What does the fine-tuning-with-trl skill do?

Fine-tune and align language models using reinforcement learning with TRL, including SFT, DPO, PPO, GRPO, and reward model training. It is listed under AI & ML, Model Training & Fine-tuning on SkillMD.

Is fine-tuning-with-trl safe to use?

SkillMD's automated safety review verdict for this skill is PASS. Independent scanners report: SkillSpector: PASS, Skill Scanner: PASS. Capability flags: makes network calls. SkillMD never runs a skill's scripts for you; review the SKILL.md before installing.

Which AI agents work with fine-tuning-with-trl?

This skill is tagged as working with Claude Code, Claude.ai, OpenAI Codex. SKILL.md is an open format, so most agents that read a skills directory can load it too.

Is fine-tuning-with-trl free to use?

Yes. Installing skills from SkillMD is free. This skill is licensed under MIT.

Who published fine-tuning-with-trl?

Orchestra Research (@orchestra-research) published this skill. Their other Agent Skills are listed on their SkillMD profile.