Fine Tuning With Trl

Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training. Use when need RLHF, align model with preferences, or train from human feedback. Works with HuggingFace Transformers.

nota-america Updated

File contents

nota-america/forgecat-agent-profiles/tree/main/profiles/orchestra-research/ai-research-skills/for-codex/.agents/skills/ai-research-skills/06-post-training/trl-fine-tuning commit e059804fcd

Frequently asked questions

npx skillmds@latest add nota-america/fine-tuning-with-trl