grpo-rl-training

orchestra-research/grpo-rl-training · Agent Skill (multi-file)

by Orchestra Research · bundle

Published · Last updated


Expert guidance for implementing GRPO/RL fine-tuning with TRL for reasoning and task-specific model training.

SKILL.md

Files

This skill is a package of 4 files. Install with the command above, or download the folder.

  • 📄SKILL.md entry
  • 📁examples
  • ⚙️reward_functions_library.py 11.3 KB
  • 📁templates
  • ⚙️basic_grpo_training.py 6.0 KB
  • 📄README.md 3.4 KB

Related

  1. trl-training · huggingface
    Train and fine-tune transformer language models using TRL (Transformers Reinforcement Learning) with support for SFT, DPO, GRPO, KTO, RLOO, and reward model training via CLI commands.
    10.8k
    repo stars
  2. fine-tuning-with-trl · orchestra-research bundle
    Fine-tune and align language models using reinforcement learning with TRL, including SFT, DPO, PPO, GRPO, and reward model training.
    10.4k
    repo stars
  3. train-sentence-transformers · huggingface bundle
    Train or fine-tune sentence-transformers models for retrieval, similarity, clustering, classification, and reranking, with support for bi-encoders, cross-encoders, and sparse encoders.
    10.8k
    repo stars
  4. huggingface-llm-trainer · huggingface bundle
    Train or fine-tune language and vision models using TRL or Unsloth on Hugging Face Jobs cloud infrastructure, with support for SFT, DPO, GRPO, and reward modeling, plus GGUF conversion for local deployment.
    10.8k
    repo stars
  5. tao-finetune-cosmos-embed · nvidia bundle
    Fine-tune, evaluate, run inference, and export Cosmos-Embed1 video-text embedding models for tasks like text-to-video retrieval and semantic deduplication.
    2.2k
    repo stars
  6. transformers · k-dense-ai bundle
    Load pre-trained models from Hugging Face Hub, run pipeline inference, generate text, and fine-tune models on NLP, vision, audio, and multimodal tasks using the Transformers library.
    30.2k
    repo stars

Frequently asked questions

How do I install the grpo-rl-training skill?

Run npx skillmds add orchestra-research/grpo-rl-training in your terminal (requires Node.js), paste this page's agent-chat prompt into Claude, Cursor, or any MCP-connected agent, or download the SKILL.md file and copy it into your agent's skills directory.

What does the grpo-rl-training skill do?

Expert guidance for implementing GRPO/RL fine-tuning with TRL for reasoning and task-specific model training. It is listed under AI & ML, Model Training & Fine-tuning on SkillMD.

Is grpo-rl-training safe to use?

SkillMD's automated safety review verdict for this skill is PASS. Independent scanners report: SkillSpector: PASS, Skill Scanner: PASS. Capability flags: executes scripts, makes network calls. SkillMD never runs a skill's scripts for you; review the SKILL.md before installing.

Which AI agents work with grpo-rl-training?

This skill is tagged as working with Claude Code, Claude.ai, OpenAI Codex. SKILL.md is an open format, so most agents that read a skills directory can load it too.

Is grpo-rl-training free to use?

Yes. Installing skills from SkillMD is free. This skill is licensed under MIT.

Who published grpo-rl-training?

Orchestra Research (@orchestra-research) published this skill. Their other Agent Skills are listed on their SkillMD profile.